The Surer a Model Is, the Less It Hears "Not"

A new Seoul National University paper finds LLMs repeat the original answer under negation 37 to 71% of the time, and more often when they're confident. The fix that looks best on standard metrics quietly teaches one stock wrong answer.

Share
The Surer a Model Is, the Less It Hears "Not"
Photo by Abdularhman Khewani on Unsplash

Ask a language model "What is the capital of Spain?" and you get Madrid. Ask "What is not the capital of Spain?" and, depending on the model and the benchmark, there is a good chance you get Madrid again.

That failure is old news. Negation has embarrassed language models since BERT. What caught my attention in a new paper from Seoul National University by Jongwook Yoon, Jongwon Lim, and colleagues is a detail buried in the second section: the failure gets more frequent as the model gets more confident in its original answer. Knowing the answer well should make it easier to rule out. In these models it makes ruling it out harder.

The numbers

The authors built a benchmark of 105,048 paired prompts from six sources (PopQA and RippleEdits for facts, PhantomWiki and SynthWorlds for logical reasoning, GQA and PTR for visual questions). Each question gets four negated versions using "not," "n't," "never," and "by no means." Then they asked a simple thing: among cases where a model got the original question right, how often did it change its answer once the "not" went in?

The answer change rates in Table 1 run from 28.9% for Llama 3.1-8B-Instruct to 62.6% for Gemma 3-27B-IT. The closed models do not escape. GPT-5.6 Luna changes its answer 48.0% of the time and Claude Sonnet 5 60.3%. Flip those around and you get the abstract's headline: in 37 to 71% of originally correct cases, the model repeats itself.

When a model does change its answer, it almost always stays on topic. Category preservation sits between 87.2% and 97.8%. Ask about a capital and you get another city, not a vegetable. So the model clearly registers that something about the question changed. It just often fails to move far enough.

How the model actually does "not"

A white and red no smoking sign on a building wall
Photo by Tarik Haiga on Unsplash

The mechanistic half of the paper, run mostly on Gemma 3-12B-IT, describes negation as three moves happening at once.

  1. Some middle-layer attention heads pay less attention to the tokens that retrieve the original answer. One head the authors single out, L31H13, sees its attention to answer tokens fall from 36.3% to 8.6% in successful cases.
  2. Other heads pay more attention to the category part of the prompt ("the capital of"), keeping the answer a city.
  3. A set of MLP neurons pushes particular candidates within that category.

The third piece is the strange one. Those neurons favor similar candidates within a category even when the original answer differs. Nothing in that circuit computes "any city except Madrid." Madrid gets turned down and a stock alternative gets turned up. The authors call that stock pick negation bias, and it explains the second failure mode in the paper: when the model's favorite negated answer happens to be the correct original answer, the answer change rate drops sharply.

The authors contrast this with research on how people process negation, where the original meaning gets retrieved first and then used to decide what to exclude. In the model, the original answer mostly gets quieted, and the replacement comes off a preference list that does not adjust much to what was negated.

Why confidence hurts

Put those pieces together and the confidence result stops looking paradoxical. Negation in these models works like a subtraction of roughly fixed size. If the model only weakly believed Madrid, a fixed push is enough to knock it off the top. If the model strongly believes Madrid, the same push leaves Madrid in first place.

The failure analysis supports that reading. In repeated-answer cases the circuitry still fires, just more weakly: L31H13's attention to the answer drops from 38.1% to only 17.2%. Doubling the selected neurons' negation-induced changes cuts the rate at which the model repeats its original answer from 100% to 58.6% on those failures. The mechanism is present. It is underpowered for exactly the questions where the model is most sure.

That inversion is a little uncomfortable to think about. The facts a model knows best, the ones you would trust it on, are the ones where "not" is most likely to bounce off.

The fix that teaches a favorite wrong answer

The part I would point practitioners at is Table 2, where the authors compare fixes on Gemma 3-12B-IT trained on the PopQA split.

Their method, Adaptive Logit Inversion Training (ALiT), takes the model's own negation-induced shift in logits and extrapolates it until some alternative beats the original answer by a margin that scales with how confident the original prediction was. Then it trains the model toward that target. Answer change goes from 58.1% to 96.6% in-domain and from 46.9% to 85.0% on the five out-of-domain benchmarks, while category preservation stays above 97% and the general capability average moves from 78.04 to 77.68.

The baselines are more instructive. Unlikelihood training and DPO push answer change to nearly 100%, but category preservation collapses to 54.0% and 42.8% in-domain. The model learns to say something other than Madrid, and often that something is no longer a city.

Plain supervised fine-tuning looks great on the first two columns. Answer change hits 96.8%, category preservation hits 100%. Then you look at answer diversity, which measures how much the model's negated answers within a category spread out instead of piling onto one. It falls from 90.9% to 8.0%. SFT taught the model to answer every "what is not the capital of X" with roughly the same city.

That is negation bias turned into policy. The paper shows the cost directly. When the original answer is different from the model's most frequent negated answer, SFT lifts answer change from 53.2% to 94.8%. When the two match, SFT drags it down from 34.2% to 18.8%. ALiT raises both groups, to 67.1% and 40.2%. Still far from solved on the hard group, but moving the right way.

If you were scoring these methods on answer change and category preservation alone, which are the two metrics most negation benchmarks report, SFT would look like a winner. You would ship a model that has memorized a default wrong answer and fails more often on exactly the case where the default collides with the truth.

What I take from it

Some caveats first. The mechanism story comes from open models in the 4B to 12B range, with Gemma 3-12B-IT doing most of the work. The closed models appear only in the behavioral evaluation. The prompts are templated questions with one fill-in answer, which is a long way from negation inside a contract clause or a medical instruction. Benchmark validation and answer grading both lean on an LLM judge, though the authors report human agreement checks. The code and data are promised on publication and not out yet.

Still, two points seem likely to travel beyond this setup.

First, a model's confidence is a poor guide to how well it will handle a twist on a question. Here the most confident answers were the stickiest, and that is the opposite of what a calibration-minded reader would hope for.

Second, an evaluation for a linguistic skill needs a metric that catches shortcut solutions. Answer diversity is a crude one, a single number about how concentrated the outputs are, and it was the only column in Table 2 that exposed SFT. Plenty of capability evals report something like accuracy and validity and stop there. This paper is a small worked example of a fix that passes both and is wrong in a systematic way.

I keep coming back to the Paris example in the paper's Figure 1. The model gets "not the capital of Spain" right by reaching for Paris. The paper's own analysis traces that reach to a category favorite, and a favorite works only until someone asks about the favorite.

0 subscribers
0 average monthly readers