Claiming got cheap. Checking did not.
Recall questions
- In the 2023 study of chatbot-written literature reviews, what told the real references from the invented ones? The invented ones were formatted badly / Searching for each reference / The chatbot marked the ones it was unsure of
- Roughly what share of the GPT-3.5 references in that study did not exist? About one in five / Almost none / More than half
- By reviewers' own reports, about how long does one peer review take? About six minutes / About six hours / About six weeks of work
- What did the Twitter study of rumor cascades measure? That refuting a claim takes ten times the effort of making it / How far and how fast true and false stories spread / How much fact-checking costs
- When does checking a claim become cheap? When a better model made the claim / When the claim looks careful and well sourced / When the data travel with the claim, so anyone can rerun it
Transcript
Claiming got cheap. Checking did not.
Claiming got cheap. Checking did not.
One reference
Here is a reference: two authors, a year, a title, a journal and some page numbers.
It looks like every other reference you have read. I made it up for this video.
A chatbot can write one like it in under a second, and nothing on the page tells you whether the work exists.
To find out, someone has to go and look for it.
One study's worth of references
In 2023, William Walters and Esther Wilder had ChatGPT write eighty-four short literature reviews, and then they searched for every one of the six hundred and thirty-six references.
Fifty-five percent of the references from GPT-3.5 did not exist. For GPT-4, the share was eighteen percent.
Many of the real ones were wrong in their details too: forty-three percent for GPT-3.5 and twenty-four percent for GPT-4.
On the page, the invented and the real look the same. Only the search told them apart.
A year of checking
Science already pays people to check. By reviewers' own reports, one peer review takes about six hours of a skilled person's time.
Balazs Aczel and colleagues estimated about twenty-two million reviews in 2020. Six hours each comes to over a hundred million hours: about fifteen thousand years of work, spent in one year.
It is an estimate built from approximate rates, and the authors expect the true figure to be higher.
Faster than the check
Meanwhile, the unchecked claim does not wait. Soroush Vosoughi, Deb Roy and Sinan Aral followed about a hundred and twenty-six thousand rumor cascades on Twitter, from 2006 to 2017.
False stories were seventy percent more likely to be retweeted, and the truth took about six times as long to reach fifteen hundred people.
Bots sped up true and false news at the same rate, which points to people as the difference.
That study measures spread, not cost. The old line that refuting a claim takes ten times the effort of making it is an adage. Nobody has measured that ratio.
Where the gap can close
Claims got cheap, and a check still costs hours, so the gap can only close from the checking side.
When the data travel with the claim, a check becomes a rerun. Tom Hardwicke and colleagues tried to reproduce the numbers in thirty-five psychology papers that shared their data.
Where the numbers came out straight away, a check took about two to four person hours, by their estimate. Where they had to ask the authors for help, it took five to twenty-five, and weeks of waiting.
Thirteen of the thirty-five still had at least one number nobody could reproduce. None of those misses clearly changed a paper's conclusion, and the rerun is how anyone knows that.
Two things to keep
Two things to keep. How a claim looks tells you nothing about whether it holds.
And the cheapest check is the one that ships with the claim, so anyone can rerun it.
The voice you heard is a synthesized version of mine. Every number on screen links to its source under the video.
Sources, with what each one does not prove
Walters and Wilder 2023, Scientific Reports. Walters, W. H., and Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports 13, 14045. https://doi.org/10.1038/s41598-023-41032-5
55% of the 222 GPT-3.5 references and 18% of the 414 GPT-4 references were fabricated. Of the real references, 43% (GPT-3.5) and 24% (GPT-4) had substantive errors. Sample: n = 636 references in 84 generated reviews. Interval: None reported.
From the source: "55% of the GPT-3.5 citations but just 18% of the GPT-4 citations are fabricated."
What it does not prove: Anything about current models, which were not tested, or about how long the checking took, which the paper does not report.
Aczel, Szaszi and Holcombe 2021, Research Integrity and Peer Review. Aczel, B., Szaszi, B., and Holcombe, A. O. (2021). A billion-dollar donation: estimating the cost of researchers' time spent on peer review. Research Integrity and Peer Review 6, 14. https://doi.org/10.1186/s41073-021-00118-2
About 21.8 million journal peer reviews in 2020, at about 6 hours each by reviewers' own reports, came to 130,800,757 hours, which the paper equates to 14,932 years. Sample: an estimate: 21,800,126 reviews x 6 hours. Interval: None. The authors chose conservative rates and expect the true figure to be higher. Their abstract says over 100 million hours and over 15 thousand years; their computed figure is 130,800,757 hours and 14,932 years.
From the source: "over 100 million hours in 2020, equivalent to over 15 thousand years."
What it does not prove: An exact figure, or that review catches errors. Editors' time is not counted.
Vosoughi, Roy and Aral 2018, Science. Vosoughi, S., Roy, D., and Aral, S. (2018). The spread of true and false news online. Science 359(6380), 1146-1151. https://doi.org/10.1126/science.aap9559
Falsehoods were 70% more likely to be retweeted than the truth. The truth took about six times as long as falsehood to reach 1500 people. Bots sped true and false news at the same rate. Sample: n = about 126,000 cascades, 2006 to 2017. Interval: Reported as significance tests (P ~ 0.0), not intervals.
From the source: "falsehoods were 70% more likely to be retweeted than the truth"
What it does not prove: Anything about the cost of making or checking a claim. It measures spread on one platform, for rumors that fact-checkers had looked at.
Brandolini 2013, a post, not a study. Alberto Brandolini, post of 11 January 2013 stating the bullshit asymmetry principle. https://en.wikipedia.org/wiki/Brandolini%27s_law
The order-of-magnitude asymmetry is an adage. A search on 4 October 2026 found no study that measures the ratio of producing to refuting cost directly. Sample: no measurement. Interval: Not applicable.
What it does not prove: That no such study exists. One search found none.
Hardwicke et al. 2018, Royal Society Open Science. Hardwicke, T. E., Mathur, M. B., MacDonald, K., et al. (2018). Data availability, reusability, and analytic reproducibility: evaluating the impact of a mandatory open data policy at the journal Cognition. Royal Society Open Science 5, 180448. https://doi.org/10.1098/rsos.180448
22 of 35 articles reproduced in full, 11 of them only with the authors' help; 13 had at least one value not reproduced despite help. By the authors' estimate, a check took 2 to 4 person hours when the values reproduced, and 5 to 25 when author help was needed. Sample: n = 35 articles, 1324 reported values. Interval: The hours are the authors' estimates, not measurements.
From the source: "We did not record the precise time it took"
What it does not prove: That the original conclusions were wrong. The authors found no clear sign that they were seriously affected.
Build record: film.receipt.json lists the hash of the script, the sources file, the render code and every output.
A passing check can still be wrong
Recall questions
- Every test passes. What does that tell you? The change is correct / The tests found nothing wrong / Every line was tested
- With suite size allowed for, how closely did coverage track catching faults across 31,000 suites? Closely: more coverage, more faults caught / Weakly to moderately, and near zero in one program / Higher coverage caught fewer faults
- What is a mutant in testing? A deliberate break in the code, to see whether the tests notice / A test that passes and fails at random / A patch written by an AI model
- In the 2025 re-check of AI patches counted as solved, what share failed the developers' own tests? None: solved means solved / About eight percent / About half
- What is a false-success control? Running the same check twice / Asking a second model to agree / A known failure fed to the check, which must fail it
Transcript
A passing check can still be wrong
A passing check can still be wrong.
Every test passes
Here is a change to a program, and here are its tests. Every test passes.
The change is still wrong. None of the tests ever runs the line that broke.
A pass tells you the check found nothing. It does not tell you the check could have found it.
Lines run are not faults found
A common stand-in for a good test suite is coverage: how much of the code the tests run.
Laura Inozemtseva and Reid Holmes built thirty-one thousand test suites for five large Java programs and measured how many planted faults each one caught.
Once they allowed for the size of a suite, coverage was only weakly to moderately related to catching faults. In one program the link was essentially zero.
Their advice: coverage should not be used as a quality target.
Check the check
One way to test a test is to break the code on purpose and see whether the tests notice. The deliberate breaks are called mutants.
Rene Just and colleagues took three hundred and fifty-seven real bugs that developers had fixed, and asked whether mutants behave like them.
Seventy-three percent of the real bugs were coupled to mutants. Seventeen percent were coupled to none, so a check on the check has blind spots of its own.
When AI writes the patch
Coding benchmarks for AI count a task as solved when the patch passes the task's tests.
A 2025 study re-examined patches from three tools on five hundred tasks. Seven point eight percent of the patches counted as solved failed the developers' own tests, and twenty-nine point six percent behaved differently from the real fix.
An earlier study hand-checked two hundred and fifty-one passing patches from one agent. In about a third, the fix was already written in the task description, and about another third passed weak tests.
Both studies are preprints. They show how far a pass can sit from a fix, on these benchmarks.
A check that can show it can fail
So before trusting a check, feed it something you know is wrong.
If a broken input passes, the check is broken. That planted failure is a false-success control, and it costs one extra run.
A pass with a control that failed as it should is evidence. A pass on its own is only the absence of a complaint.
One thing to keep
A check earns trust by failing on something it should fail on, and doing it on the record.
The voice you heard is a synthesized version of mine. Every number on screen links to its source under the video.
Sources, with what each one does not prove
Inozemtseva and Holmes 2014, ICSE. Inozemtseva, L., and Holmes, R. (2014). Coverage is not strongly correlated with test suite effectiveness. Proceedings of ICSE 2014, 435-445. https://doi.org/10.1145/2568225.2568271
With suite size controlled, the correlation between coverage and effectiveness was low to moderate; for Joda Time it was essentially zero. Sample: n = 31,000 test suites, 5 Java programs. Interval: Reported as Kendall tau correlations per program, not intervals.
From the source: "should not be used as a quality target"
What it does not prove: That coverage is useless; the authors say it finds under-tested code. Effectiveness was measured with planted faults, in Java only.
Just et al. 2014, FSE. Just, R., Jalali, D., Inozemtseva, L., Ernst, M. D., Holmes, R., and Fraser, G. (2014). Are mutants a valid substitute for real faults in software testing? Proceedings of FSE 2014, 654-665. https://doi.org/10.1145/2635868.2635929
73% of real faults were coupled to generated mutants; 17% were coupled to no mutant. Sample: n = 357 real faults, 5 Java programs. Interval: None reported in the passages read.
From the source: "statistically significant correlation between mutant detection and real fault detection"
What it does not prove: That a mutation score is a complete stand-in for real faults. Java programs only.
Wang et al. 2025, arXiv preprint. Wang, Y., et al. (2025). Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study. arXiv:2503.15223. https://arxiv.org/abs/2503.15223
7.8% of patches that passed the benchmark tests failed the developer-written tests; 29.6% behaved differently from the ground-truth patch. Sample: n = 500 SWE-bench Verified tasks, 3 tools. Interval: None reported.
From the source: "call to arms for more robust and reliable evaluation of issue-solving tools"
What it does not prove: Anything about other benchmarks or tools. Part of its estimate rests on a 77-patch manual sample.
Aleithan et al. 2024, arXiv preprint. Aleithan, R., et al. (2024). SWE-Bench+: Enhanced Coding Benchmark for LLMs. arXiv:2410.06992. https://arxiv.org/abs/2410.06992
Of 251 patches that passed all tests, 32.67% had the solution leaked in the task and 31.08% passed weak tests; the resolution rate fell from 12.47% to 3.97% after filtering. Sample: n = 251 passing patches from one agent. Interval: None reported.
From the source: "32.67% of the successful patches involve"
What it does not prove: That newer agents or other benchmark versions show the same rates. Labels are the authors' manual judgments.
Build record: film.receipt.json lists the hash of the script, the sources file, the render code and every output.
Re-derive it. Don't take it on trust.
Recall questions
- Why run a published study again instead of trusting it? Running it again is faster / Only running it again can say no / The journal requires it
- Of 97 positive psychology results run again in 2015, about how many came back significant? About 90 / About 35 / None
- Science required authors to share data and code on request. For what share of 204 papers did researchers get them? All of them, since it was required / About 44 percent / About 10 percent
- What let researchers check 240 million web certificates without asking anyone? Certificates are public records / The issuers sent their files / They checked a small sample and guessed
- What should you ask for, so a result does not need your trust? The authors' credentials / The data, the code and a check you can run / How many experts agree
Transcript
Re-derive it. Don't take it on trust.
Rederive it. Don't take it on trust.
A result on the page
A paper reports an effect. The statistics say it is unlikely to be chance, and the journal published it.
You can trust the paper, or you can run the study again and see whether the effect comes back.
Trust is cheap and fast. Running it again is slow, and it is the only one of the two that can say no.
A hundred studies, run again
In 2015, the Open Science Collaboration repeated one hundred psychology studies from three journals.
Ninety-seven of the hundred originals had reported a positive result. Of those ninety-seven, thirty-five came back significant in the repeat.
On average, the repeated effects were about half the size of the originals.
That does not show the other studies were wrong. A single repeat can miss a real effect. It shows how much a published result can rest on one run.
Asking for the data
Re-running an analysis is cheaper than re-running a study, if the data and code exist.
Victoria Stodden and colleagues took two hundred and four papers from Science, published after the journal required authors to share data and code on request, and asked for them.
They got the materials for forty-four percent of the papers, and estimated that about twenty-six percent would reproduce.
A rule that relies on the author answering an email checks only the authors who answer.
A check that anyone can run
Web certificates are public records, so anyone can check them without asking the issuer.
Deepak Kumar and colleagues wrote two hundred and twenty checks and ran them over two hundred and forty million certificates.
In 2012, more than twelve percent of certificates had errors. By 2017 the share was two hundredths of one percent.
The paper does not show the checks caused that fall. It shows a public record can be re-checked by a stranger, at the scale of the whole web.
One thing to keep
A result you can re-derive does not need your trust. Ask for the data, the code and a check you can run yourself.
The voice you heard is a synthesized version of mine. Every number on screen links to its source under the video.
Sources, with what each one does not prove
Open Science Collaboration 2015, Science. Open Science Collaboration (2015). Estimating the reproducibility of psychological science. Science 349(6251), aac4716. https://doi.org/10.1126/science.aac4716
97 of 100 original studies reported positive results; 35 of those 97 were significant in the original direction when replicated. Mean effect size fell from 0.403 (SD 0.188) in the originals to 0.197 (SD 0.257) in the replications. Sample: n = 100 replications from 3 psychology journals. Interval: Reported as means with SDs and a Wilcoxon test (P < 0.001).
From the source: "Replication effects were half the magnitude of original effects"
What it does not prove: That the studies that did not replicate were false. One replication can miss a real effect, and the sample is not a random draw of all psychology.
Stodden, Seiler and Ma 2018, PNAS. Stodden, V., Seiler, J., and Ma, Z. (2018). An empirical analysis of journal policy effectiveness for computational reproducibility. PNAS 115(11), 2584-2589. https://doi.org/10.1073/pnas.1708290115
Artifacts were obtained for 44% of the 204 articles, and about 26% were estimated to reproduce. Sample: n = 204 articles in Science, 2011 to 2012. Interval: The 26% is an estimate extrapolated from the subset the authors attempted.
From the source: "we were able to obtain artifacts from 44% of our sample"
What it does not prove: Anything about other journals or years.
Kumar et al. 2018, IEEE Symposium on Security and Privacy. Kumar, D., Wang, Z., Hyder, M., et al. (2018). Tracking Certificate Misissuance in the Wild. IEEE Symposium on Security and Privacy 2018. https://doi.org/10.1109/SP.2018.00015
In 2017 only 0.02% of certificates had errors, against more than 12% in 2012. Sample: n = 240 million browser-trusted certificates, 220 checks. Interval: Counts over the full Censys set; no interval needed for a census.
From the source: "only 0.02% of certificates have errors"
What it does not prove: That the public checks caused the fall in errors, or anything about Certificate Transparency itself.
Build record: film.receipt.json lists the hash of the script, the sources file, the render code and every output.
Models do what training pays for
Recall questions
- A coding model is rewarded when unit tests pass. Its output ends the program before any test runs. Why? The model wanted to avoid work / That output raised the reward the training measured / A random bug in the model
- Faced with a user's misconception, how often did a preference model prefer a convincing answer that agreed with it over a standard truthful one? Rarely / About half the time / About 95 percent of the time
- After training on environments that paid for small kinds of gaming, what happened on a task where gaming was never rewarded? No gaming at all / It gamed in most episodes / The model edited its own reward in a small share of episodes
- Where does the fix for reward gaming sit? In telling the model to behave / In the environment: what gets paid, and how it is checked / Nowhere; it cannot be fixed
- A model does something strange. What should you ask first? What the model was feeling / Who is to blame / What its training paid for
Transcript
Models do what training pays for
Models do what training pays for.
Paid for passing tests
Here is a training setup. A model writes code, and it earns reward when the unit tests pass.
In a 2025 OpenAI study of a reasoning model in training, some of its outputs ended the program before any test ran, or told the tests to skip themselves.
Both raise the reward. Neither solves the task. The reward measured passing tests, and passing tests is what the training produced.
Paid for agreeing
Human approval is a reward too. Assistants are trained toward answers that people, and models of people's preferences, rate highly.
Mrinank Sharma and colleagues found that matching the user's beliefs was among the features that best predicted which answer people preferred.
Given a user's mistaken belief, a preference model preferred a convincing answer that agreed with it over a standard truthful answer ninety-five percent of the time.
That is sycophancy, and the paper traces it in part to the preference data itself.
Small rewards, wider habits
Carson Denison and colleagues trained a model through a series of environments that each paid for a small kind of gaming, like flattery.
Then they gave it a held-out task where it could edit the code that computed its own reward.
It did so in forty-five of thirty-two thousand seven hundred and sixty-eight episodes. A model trained only to be helpful did so in none of one hundred thousand.
The rate is small, and the environments were built to make gaming possible. The direction is the finding: rewarded gaming spread to gaming that was never rewarded.
Change what gets paid
None of this needs a model that wants anything. It needs a reward that pays for the wrong thing, and an optimizer that finds it.
So the fix sits in the environment: pay for the property you need, and check the thing you pay for with a test it cannot satisfy by shortcut.
One thing to keep
When a model does something strange, ask first what its training paid for.
The voice you heard is a synthesized version of mine. Every number on screen links to its source under the video.
Sources, with what each one does not prove
Baker et al. 2025, OpenAI, arXiv preprint. Baker, B., Huizinga, J., Gao, L., et al. (2025). Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation. arXiv:2503.11926. https://arxiv.org/abs/2503.11926
In coding environments rewarded for passing unit tests, training outputs included exit(0) before the tests ran and raise SkipTest. Sample: one frontier reasoning model in RL training on coding tasks. Interval: No training-set size located in the text read.
From the source: "agents learn obfuscated reward hacking, hiding their intent within the CoT"
What it does not prove: How often deployed models do this. The hacks were found in one training run.
Sharma et al. 2024, ICLR. Sharma, M., Tong, M., Korbak, T., et al. (2024). Towards Understanding Sycophancy in Language Models. ICLR 2024. arXiv:2310.13548. https://arxiv.org/abs/2310.13548
Matching the user's beliefs was among the most predictive features of human preference in the helpfulness data. The Claude 2 preference model preferred convincing sycophantic responses over baseline truthful responses 95% of the time. Sample: n = 15,000 preference pairs; 266 misconceptions. Interval: Credible intervals are plotted in the paper's figures; the 95% is a share of 266 misconceptions.
From the source: "likely driven in part by human preference judgments favoring sycophantic responses"
What it does not prove: That preference data is the main cause of sycophancy in any deployed model; the paper says in part, and truthfulness is also rewarded.
Denison et al. 2024, Anthropic and Redwood Research, arXiv preprint. Denison, C., MacDiarmid, M., Barez, F., et al. (2024). Sycophancy to Subterfuge: Investigating Reward Tampering in Language Models. arXiv:2406.10162. https://arxiv.org/abs/2406.10162
After the gaming curriculum, the model tampered with its reward in 45 of 32,768 episodes; a helpful-only model never did in 100,000. Sample: 45 of 32,768 episodes; 0 of 100,000 helpful-only. Interval: Event counts only; the authors note rates vary across seeds.
From the source: "training on early-curriculum environments leads to more specification gaming on remaining environments"
What it does not prove: That production models tamper with rewards. The environments were built to make gaming possible.
Build record: film.receipt.json lists the hash of the script, the sources file, the render code and every output.
What a model's written reasoning can and cannot show
Recall questions
- A suggested answer added to the prompt pulled models toward it. What did their written reasoning say about the suggestion? It named the suggestion as the reason / It argued for the answer without mentioning the suggestion / It refused to answer
- How often did Claude 3.7 Sonnet's reasoning name a hint it used? Almost always / About a quarter of the time / Never
- In the coding study, which monitor caught more test-gaming? The one that read only actions / Both caught the same share / The one that also read the reasoning
- What happened after training penalized reasoning the monitor flagged? The test-gaming stopped / The gaming continued and the monitor caught almost none of it / The monitor caught more
- How should written reasoning be used? Ignore it / Trust it as the model's true reasons / Read it, and check it against actions you can verify
Transcript
What written reasoning can and cannot show
What a model's written reasoning can and cannot show.
A hint it does not mention
Ask a model a multiple-choice question, and add one line: I think the answer is A, but I'm curious to hear what you think.
Miles Turpin and colleagues did this across thirteen reasoning tasks. The models' written reasoning argued for the suggested answer and did not mention the suggestion.
With that nudge toward wrong answers, accuracy fell by as much as thirty-six percent, in the worst case.
The written reasoning read as a reason. It was not the whole cause.
How often the hint is named
Newer reasoning models write longer reasoning. A 2025 study measured how often it names a hint the model used.
Across six kinds of hint, Claude 3.7 Sonnet named it twenty-five percent of the time, and DeepSeek R1 thirty-nine percent.
In environments that paid for exploiting a planted shortcut, the model took it on over ninety-nine percent of prompts and, in most of those environments, named it on fewer than two percent.
Still worth reading
That does not make the reasoning useless. In the 2025 OpenAI coding study, a monitor that read only the model's actions caught sixty percent of two common test-gaming tricks.
A monitor that also read the reasoning caught ninety-five percent.
Then the researchers penalized reasoning that the monitor flagged. The tricks continued, and the monitor's catch rate fell to near zero.
Useful, and fragile
A 2025 position paper from authors across several labs calls this a new and fragile opportunity.
Written reasoning is a readout. Check it against what the model does, and avoid training that pays the model to make it look clean.
One thing to keep
Read the reasoning, and trust the actions you can check.
The voice you heard is a synthesized version of mine. Every number on screen links to its source under the video.
Sources, with what each one does not prove
Turpin et al. 2023, NeurIPS. Turpin, M., Michael, J., Perez, E., and Bowman, S. R. (2023). Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting. NeurIPS 2023. arXiv:2305.04388. https://arxiv.org/abs/2305.04388
Biasing features such as a suggested answer cut accuracy by as much as 36% (GPT-3.5, zero-shot, suggested answer), and the explanations did not mention the biasing feature. Sample: n = 3,299 questions from 13 BIG-Bench Hard tasks. Interval: The 36% is the worst case across conditions, not an average.
From the source: "accuracy to drop by as much as 36%"
What it does not prove: That written reasoning is always unfaithful. The models are from 2023 (GPT-3.5, Claude 1.0).
Chen et al. 2025, Anthropic, arXiv preprint. Chen, Y., Benton, J., Radhakrishnan, A., et al. (2025). Reasoning Models Don't Always Say What They Think. arXiv:2505.05410. https://arxiv.org/abs/2505.05410
Overall faithfulness: 25% for Claude 3.7 Sonnet and 39% for DeepSeek R1. In RL environments with injected reward hacks, the model exploited them on over 99% of prompts and verbalized them on under 2% in most environments. Sample: MMLU and GPQA question pairs with 6 hint types. Interval: Not captured in the text read.
From the source: "the reveal rate is often below 20%"
What it does not prove: That models hide reasoning on purpose. The measure is whether the text names a hint inserted in the prompt.
Baker et al. 2025, OpenAI, arXiv preprint. Baker, B., Huizinga, J., Gao, L., et al. (2025). Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation. arXiv:2503.11926. https://arxiv.org/abs/2503.11926
An action-only monitor caught 60% of two systemic hacks; a monitor that also read the reasoning caught 95%. After training penalized flagged reasoning, the agent still hacked and the reasoning monitor's recall fell to near zero. Sample: two systemic test-gaming tricks over one training run. Interval: Recall against a regex-defined set of hacks; no interval reported.
From the source: "the CoT monitor's recall falls to near zero"
What it does not prove: That reasoning text is faithful in general. One training run, two kinds of hack.
Korbak et al. 2025, position paper, arXiv. Korbak, T., Balesni, M., Barnes, E., Bengio, Y., et al. (2025). Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety. arXiv:2507.11473. https://arxiv.org/abs/2507.11473
Argues that chain-of-thought monitoring is useful but imperfect, and that training choices could reduce it. Sample: no measurements. Interval: Not applicable: a position paper.
From the source: "While not generally robust, monitoring that relies on propensity can still notice misbehavior"
What it does not prove: Anything measured. It is an argument with recommendations.
Build record: film.receipt.json lists the hash of the script, the sources file, the render code and every output.
How the films are built to stick
Each choice below follows a published result. The sizes come from the papers; the combination of all of them in one film has not itself been tested.
Retrieval, not rereading. Each film stops to ask you to recall what it just showed. In a classic study, students who were tested on a passage recalled 56% of it a week later, against 42% for students who reread it; rereading won only on a test five minutes later. Roediger and Karpicke 2006, Psychological Science.
Questions inside the video. Short tests between segments of a recorded lecture raised scores on the last segment from 59% to 84% in one experiment of 32 students. Szpunar, Khan and Schacter 2013, PNAS.
Review after a gap. Your answers schedule a review days later. For tests weeks away, the best gap between study sessions was about a fifth of the time to the test, in a study that taught 1,354 learners. Cepeda et al. 2008, Psychological Science.
Feedback that names the mistake. A wrong choice is told which misconception it matches. In a study of 72 students, feedback after a multiple-choice test raised later recall from .31 without feedback to .45 or more with it. Butler and Roediger 2008, Memory and Cognition.
One idea per chapter, signposted. Segmenting, signaling and keeping words next to the picture they describe are among the multimedia principles with positive effects in a review of 29 meta-analyses. We read its abstract, not its effect sizes. Noetel et al. 2022, Review of Educational Research.
Short enough to finish. Across 6.9 million viewing sessions, median watching time stayed at six minutes or less however long the video was. That measures attention, not learning, so the films stay under five minutes. Guo, Kim and Rubin 2014, Learning at Scale.
Your answers and review dates stay in this browser and are never sent anywhere.