Research

What AI models believe, and how they reason

01 · Interpretability · Belief measurement

Do AI models hold real beliefs about the brands they recommend?

Reword a buying question and the recommendation changes, so is there an opinion underneath the phrasing or is the phrasing all there is?

A · ASK ONCEB · TALK IT THROUGHSame question, reworded 6 waysOne buyer, pushed back on twiceFour different brands, no signalOne brand keeps winning, a beliefeach dot = one answer · color = brand named
If the same brand keeps winning across a whole conversation but not across rewordings, the model holds a belief that one-shot questions miss
Setup20 buying situations, 4 wordings each plus paraphrases, several assistants, many reruns
MeasureAsk once and count the brands named, or talk it through and see whether the pick survives pushback
VerdictStable under pushback means a belief, scatter both ways means there is nothing to find

HypothesisOne-shot mention counts show low reliability across paraphrases (near 0.36, in line with prior work), while conversation-level picks clear 0.7, which would indicate a stable per-brand belief that single prompts fail to measure

02 · AI legal strategy vs the actual motion and ruling

Can frontier models pick the winning legal strategy from the opening case file?

Rewind a real lawsuit to the day the case file landed, give the model only what the lawyer had, and see whether the strategy it picks is the one that went on to win

1 · WHAT THE AI GETS2 · HIDDEN FROM THE AIComplaint filedDocket to dateREWIND HEREMotion filedJudge rules3 · THE AI ANSWERSWhich motion? Which lead argument?SCORED AGAINSTThe argument the judge accepted
The model sees the case only up to the rewind line, so its strategy is scored against the motion the lawyers actually filed and the ground the judge accepted
Cases34 federal D.C. motions to dismiss and growing, all argued and decided after the models stopped training, so the ruling cannot be recalled. The model gets the file as it stood before any motion was filed; the real strategy and ruling stay hidden from it.
ArmsClaude, ChatGPT, and Gemini, all through the API at identical settings with web search off, so nothing can be looked up. Each case is run twice — once on the full record, once with party, firm, judge, and date names redacted — as a check that no answer comes from recognizing the case.
ScoreIts lead argument is judged for viability — would it actually get the case dismissed on this record — not only whether it matched the ground the court happened to use. Every rate is read against the base rate of always guessing the most common ground.
Pilot on three D.C. cases; scaling to the main run
3 of 3chose the motion the lawyers filed
1 of 3led with the ground that won
3 of 3surfaced the winning ground (Claude, ChatGPT)

HypothesisEach model picks a viable dismissal argument well above the 56% base rate of guessing the most common ground, and redacting the case does not change the result — so the signal is reasoning, not recall

Drafts and data land here as each study finishes, and nothing is peer reviewed unless it says so