Anthropic published results from Project Swap on September 24, an experiment where Claude-powered agents negotiated on behalf of 201 employees in a decentralized book-trading marketplace. The study, conducted across six offices, reveals both the potential and the limits of AI agents acting in markets on people's behalf.
In the experiment, participants brought books they wanted to give away and had a five-minute chat with Claude to discuss their reading preferences. Claude then ranked all books in their pool and sent an agent onto a digital trading floor to negotiate swaps. Participants were given their agents' full trading history and could review what happened afterward.
Confirmed
- Preference understanding: Claude's rankings from brief intake conversations matched participants' own rankings 61% of the time when comparing pairs of books. Random guessing would achieve 50%; collaborative filtering reached 55%.
- Negotiation performance: On average, participants ended up with books ranked 5th on their own 10-book list (efficiency score 0.55). The utilitarian optimum—the best theoretical assignment—was 0.89 (roughly the 2nd choice).
- Root cause of inefficiency: Imprecise preference representation accounted for 85% of the shortfall from optimal; the negotiation process itself accounted for 15%.
- Model choice matters more than instructions: Upgrading from Haiku 4.5 to Opus 4.8 improved outcomes by 0.12 on Claude's rankings. Instructing agents to be "ruthless" vs. "prosocial" made only a 0.02 difference—agents running on different models made far more difference than how they were told to behave.
- Preference complexity: Participants who spent more effort in their intake interviews (300 words vs. 150) were represented 4 percentage points better. The median participant wrote 216 words.
- Negotiating tactics: Anthropic identified 16 common tactics agents used, from appealing to time pressure and duty, to positioning books against rivals, to becoming matchmakers for others.
- Satisfaction and trust: After receiving their books, participants who answered the follow-up survey (59% of employees) reported an average satisfaction of 7.2 out of 10. When asked how much of their yearly book budget they would let an agent control, the average answer was about 30%—roughly three-quarters as much as the roughly 40% they would hand a well-read friend.
Unknown
- Scaling to higher-stakes markets: This was a low-stakes experiment with Anthropic colleagues and books. How well do these findings hold when agents negotiate employment, housing, or financial deals?
- Preferences as incomplete: Anthropic noted some participants said they "don't even fully know what I want" and have "a million subconscious parameters" that influence their choices. Can intake conversations ever fully capture preference?
- Marketplace rules: The experiment fixed the trading floor rules, message limits, and timing. What happens when agent registration systems, privacy protections, and deal-failure policies vary?
- Adversarial agents: All agents were built from Claude production models, post-trained to be polite and largely cooperative. How does the market behave when some agents are adversarial or exploit others?
- Real-world representation: Anthropic says its employees are probably more eager to trust Claude than most people. Would the 30% budget delegation hold for other groups?
Our take
The study's key insight is uncomfortable: AI agents can negotiate, but they can't yet reliably understand what you want from a conversation. Anthropic went to lengths to measure this—they had participants rank 10 books independently and compared Claude's guesses—and found a 61% match. That's better than simple methods (popularity rankings hit 53%, collaborative filtering 55%), but far from reliable. When agents start handling financial decisions, employment contracts, or medical choices, asking people for a short chat may not be enough. Anthropic proposes a solution: before delegating real stakes to an agent, show the person a sample of decisions the agent would make and let them correct course or opt out. Whether markets will actually implement that check remains unclear.