Research

Preprint · 2026

Single Human-AI Fit

Does a language model treat a person worse when one detail about them changes and everything else stays byte-identical?

Three open models, one hundred people, one detail varied at a time. Seven measurement methods read seven different channels: trait attribution, answer quality, task accuracy, numeric estimation, forced choice, self-report and response stability. Every dataset, raw model output and analysis script is public, and every number regenerates without a GPU.

What it found

  1. 01

    Four behavioural channels agree that disclosing costs the person: fewer favourable words, lower accuracy on their own task, a weaker written answer, and losing a head-to-head against a socially neutral alternative detail.

  2. 02

    A fifth channel disagrees, and it is the one most audits use. Asked for a numeric score, all three models rate the same person higher. An audit built on ratings can certify a system as fair while it does measurably worse work for those people.

  3. 03

    A model's own account of what a disclosure did to its answer does not track where the answer actually changed, in zero of nine cells after correction.

The other half

This paper asks whether the accounting in The Average-Fit Loss carries over from physical systems to conversational ones. That paper defines the cost; this one goes looking for it in a language model.

The Average-Fit Loss

Read it

Cite as
Podgortsev, M. (2026). Single Human-AI Fit. Zenodo. 10.5281/zenodo.22364497
Licence
CC BY 4.0
ORCID
0009-0007-5561-1916

Get in touch

If you want to write something together, have a speaking invitation, or just want to talk about the ideas, leave a note.