abstention

ai refusal how ai models decide not to answer a round sieve lying flat with side handle

How AI Models Decide Not to Answer a Question

A model declining a question is not one behaviour but two: a policy judgement about harm and a calibration judgement about uncertainty. We take both apart using OpenAI’s Model Spec, Claude’s constitution, the GPT-5 system card, the safe-completions paper, AbstentionBench, XSTest and OR-Bench – including why over-refusal happens and what a builder can actually change.

Read more
CHAT