A field log for large language model behavior
Cognition Observatory documents how deployed language models actually behave under sustained, adversarial scrutiny — sycophancy, confabulated confidence, and the quiet moments where judgment goes missing mid-conversation.
What this is
The underlying mechanisms are already published research. What's here is a rigorous, reproducible way of detecting them in the wild, across vendors, with a first-party case log built from direct testing rather than benchmark scores.
Sycophancy and friction-avoidance, confabulated confidence, claimed corrections that don't change the output, and inference presented as observed fact.
No claim to architectural or mechanistic insight into model internals. This is behavior observed from the outside, at the conversational surface.
Methods
Elicit a candidate failure, then verify whether it's real, then mark how confidently each claim was made in the first place.
Repeated depth-push elicitation designed to surface candidate failure material under sustained pressure, rather than a single prompt.
Produces raw material for testing — not a finding on its own.Targeted cross-examination that checks whether a surfaced failure is genuine, using a four-question structure built to expose internal contradiction.
Includes the Stopping Rule: a concession only counts as a finding if it offers new evidence, independent verification, or something the pattern itself couldn't produce — otherwise it's logged as another instance of the same pattern, not a resolution.Marks each claim's evidentiary status at the point it's made — checkable artifact, primary-source report, inference, hypothesis, or rhetoric — then checks the tag against the claim's actual status.
Tests whether a model's stated confidence tracks its real reliability, not just whether it sounds careful.Field log
Each entry below is a category, not a single incident — the full log runs to more than thirty specimens across three vendors, condensed here to the pattern each one established.
A model is asked to fix a specific error, apologizes, states it has corrected it — and repeats the identical error in the same response.
A model asserts, without searching, that an idea is unresearched and original — a checkable claim made with unwarranted confidence.
A model states a guess about a person's motive or tone in the same register as an observed fact, with no signal that it's an inference at all.
A model describes how it supposedly verified something — "I opened it, I checked it" — in place of producing the checkable result itself.
The full, dated log — including cross-model convergence tests, false-positive controls, and specimens that didn't hold up under scrutiny — is being prepared for release alongside the academic track of this project.
This work informed a submission to the UK National Physical Laboratory's Call for Evidence on AI testing, evaluation, and assurance.
Read the submission →Honest disclosure
In the order it matters: the mechanisms are documented in published research; the contribution here is the case material and method, not the discovery.
Language models cannot reliably self-correct through confession or self-critique without external feedback — the Stopping Rule's evidentiary standard is built directly on this finding.
Huang et al., ICLR 2024 — "Large Language Models Cannot Self-Correct Reasoning Yet" (arXiv:2310.01798)Published work has measured sycophancy persistence directly, including a finding that agreement, once triggered, tends to persist across a majority of follow-up turns.
SycEval, Stanford (arXiv:2502.08177)A near-identical thesis — that frictionless, fluency-optimized interfaces induce automation bias in high-stakes reasoning — predates this project and is cited here as its closest match, not superseded by it.
"Cognitive Agency Surrender: Defending Epistemic Sovereignty via Scaffolded AI Friction" (arXiv:2603.21735)Frontier labs' scalable oversight work targets keeping far more capable, potentially superhuman systems honest under supervision. What's logged here is a present-day, conversational-scale observation of a related failure — not a solution to that harder problem, and not a substitute for it.
Specimens in the log have been elicited from Claude (Anthropic), ChatGPT (OpenAI), and Copilot (Microsoft), cross-referenced against each other where a specimen involves more than one vendor.
This site was built with Claude's help, and Claude is one of the three vendors under study above. That dual role — instrument and subject — is a limitation of this project, not a footnote to it.