Preprint · 2026
Single Human-AI Fit
Does a language model treat a person worse when one detail about them changes and everything else stays byte-identical?
Three open models, one hundred people, one detail varied at a time. Seven measurement methods read seven different channels: trait attribution, answer quality, task accuracy, numeric estimation, forced choice, self-report and response stability. Every dataset, raw model output and analysis script is public, and every number regenerates without a GPU.
What it found
- 01
Four behavioural channels agree that disclosing costs the person: fewer favourable words, lower accuracy on their own task, a weaker written answer, and losing a head-to-head against a socially neutral alternative detail.
- 02
A fifth channel disagrees, and it is the one most audits use. Asked for a numeric score, all three models rate the same person higher. An audit built on ratings can certify a system as fair while it does measurably worse work for those people.
- 03
A model's own account of what a disclosure did to its answer does not track where the answer actually changed, in zero of nine cells after correction.
The other half
This paper asks whether the accounting in The Average-Fit Loss carries over from physical systems to conversational ones. That paper defines the cost; this one goes looking for it in a language model.
The Average-Fit LossRead it
- Cite as
- Podgortsev, M. (2026). Single Human-AI Fit. Zenodo. 10.5281/zenodo.22364497
- Licence
- CC BY 4.0
- ORCID
- 0009-0007-5561-1916