Project Swap's Bottleneck Was Preference
Anthropic's Project Swap put Claude agents into a book barter market. Live efficiency hit ~0.55 vs a ~0.89 optimum; ~85% of the shortfall was preference representation, not the trading floor.
Anthropic's Project Swap (Hitzig et al., 24 Sep 2026) put 201 employees into a barter market for summer books. Each person had a short Claude intake, then a Claude agent took that preference ranking onto a decentralized trading floor and swapped. Live efficiency on people's own rankings averaged about 0.55 against a utilitarian optimum near 0.89. Roughly 85% of that shortfall traces to preference representation. The agents did fine at haggling. Knowing people was the expensive miss.
What the intake actually bought
Claude never saw participants' ground-truth rankings. From a five-minute chat (median participant: 216 words across eight messages), it ranked every book in the local pool. Against the 10-book lists people later submitted, Claude's ordering agreed on 61% of pairs. Random would be 50%. Popularity by Open Library want-to-read counts hit about 53%. A collaborative-filter baseline from co-occurrence data got about 55%.
Sixty-one percent sounds modest until you notice how little text produced it. Writing roughly 300 words instead of 150 predicted about four more points of pairwise agreement. Representation moved with effort, even inside a short conversation.

Where the efficiency leaked
Score each person's take-home book on their own ranking (1 = top listed book, 0 = last). Live floors averaged about 0.55, roughly a fifth-ranked book on a ten-book list. The utilitarian optimum on those same rankings was about 0.89.
Anthropic split the shortfall. Compute the best possible assignment from Claude's rankings, then score it on ground truth: about 0.60. Most of the distance from 0.89 to 0.55 sits in that step. About 15% sits in the decentralized floor itself. Once rankings are noisy, Top Trading Cycles from Claude's lists also lands near 0.60. Market design barely moves the needle when the preference map is blurry.
Model over instruction, once preferences are fixed
Judged on people's own rankings, agent-design tweaks look small. That is expected: every agent on every floor traded from the same Claude rankings, so model and prompt only touch the bargaining residual.
On Claude's own rankings, the picture sharpens. Model choice mattered more than ruthless versus prosocial instructions. Ruthless agents scored about 0.02 higher than prosocial ones on the same floor. Upgrading Haiku to Opus moved people about 0.12 up Claude's lists; Sonnet to Opus about 0.08.
Stronger models produced more efficient floors. With neutral instructions, Haiku floors averaged about 0.75; Opus floors about 0.88; the utilitarian optimum on Claude rankings was about 0.95. Sonnet sat between Haiku and Opus. Fable was close to Opus but a bit lower. On mixed floors, Opus halves finished ahead of weaker halves, and each half performed roughly like it did among its own kind.
What people said afterward
Endline satisfaction among respondents averaged about 7.2 out of 10. Asked what share of next year's book budget they would hand an agent that knew the intake, picked, and bought with no veto, the average was about 30%. A well-read friend who knew their taste got about 40%. Trust tracked representation quality: people who said Claude's intake recap missed nothing would give an agent about 34%; those who said something was missed would hand over about 23%.
The design implication
Project Swap is a controlled sequel to Project Deal, built so preference quality and floor outcomes can be scored separately. The agents proposed swaps, rotations, and pressure tactics; they mostly told the floor their top pick; they almost never lied about it. Negotiation was busy and competent enough. The leak was earlier.
Anyone shipping agents into markets should budget more for preference checks than for tougher negotiation prompts. Project Swap's prototype is simple: after intake, sample a few pairwise rankings and show the person how well the agent matches. That test is the product feature sitting in plain sight in the paper.
Sources: Project Swap: What happens when agents trade for us? (Hitzig, Carr, Cotter, Troy, Turman, Massenkoff, McCrory; Anthropic, 24 Sep 2026).