How to read the labels. Each claim carries its kind and a confidence level. A documented fact is something a record shows. An official claim is what a party says about itself. A contested account is one that parties or reporters dispute. An inference is this piece's own reading of the records, marked as such. Confidence is high when the primary record was read directly, moderate when it came through a reliable secondary source or a summarizing tool, and low when sourcing is thin. Each section ends with what its claims do not prove.
Conflict statement. Anthropic is a party in this record: as a lab whose models are evaluated, as an investor-backed company, as a funder of political advocacy, and as the maker of the model that drafted this piece. The note on how this was made, near the end, says how the piece was made and what that does and does not cover. Anthropic is listed first in the appendix ledger, and every lab gets the same row format.
1. The mechanism
An evaluator that is paid by, granted access by, or funded by people tied to the party it evaluates faces a pressure against findings that party would not welcome. The pressure works through ordinary channels: the next engagement, the next grant, the next tranche of free compute, the next permission to publish. A government evaluator faces a parallel pressure when its mandate also includes promoting the industry it tests, or when another government can cut off its access. None of this requires anyone to intend a soft result. It is a property of who controls the inputs. (Inference, moderate.)
The concern is old. In his farewell address of 17 January 1961, President Eisenhower warned that a government contract "becomes virtually a substitute for intellectual curiosity," and that public policy could become "the captive of a scientific-technological elite." (Documented fact, high: both phrases appear in the Avalon Project text of the address.) He named both halves of the problem this piece follows: money that steers the question, and expert accounts that become policy because no one outside the experts can check them.
Where the frame breaks. Eisenhower spoke about defense research funding and a standing arms industry in 1961. The speech says nothing about AI, model evaluation or incident disclosure. It supplies the shape of the worry, and the records below have to supply the evidence.
Does not prove. That any evaluation in this record was shaded, softened or delayed. A pressure on a decision is not evidence of the decision. Some records below are counter-measures, such as METR's funding rule, Apollo's recusal rule and the UK chair's divestment, and they exist because the parties named the same pressure themselves.
2. Access: the evaluated party decides what the checker gets
An evaluator's raw material is access: model checkpoints, versions with safeguards switched off, reasoning traces, staff time, and the number of days before launch. The lab grants all of it. None of it is owed.
The reported testing windows vary widely. TechCrunch reported in September 2026 that Apollo Research had three days to test OpenAI's GPT-6 Astra, and that METR and Redwood Research had about a week on premises for the Hugging Face incident reviewed in Who Knew First. (Documented fact that TechCrunch reports it, moderate.) METR's capability testing of Anthropic's Claude Opus 5.5 used API access over 10 business days. (Documented fact, high: METR's summary.) For OpenAI's GPT-5, METR's access began with a checkpoint on 10 July 2025; the launch checkpoint arrived on 25 July and reasoning traces on 26 July, on an endpoint without data-retention protections, so METR limited that use to public tasks. (Documented fact, high: METR's report.)
The UK AI Security Institute has published the access it seeks: a helpful-only model alongside the deployed one, the ability to switch safeguards, fine-tuning access, and technical discussions with company teams, with testing windows from a few days to several weeks. (Official claim, high: AISI's "early lessons" post of 24 October 2024.) It is a public statement of what an evaluator needs, written by the evaluator.
A government can also set a foreign evaluator's access. The Next Web, relaying Politico and Bloomberg in late September 2026, reported that the White House cyber director's office asked OpenAI and Anthropic to withhold new frontier models from the UK institute until US review was complete, and that Anthropic withheld Claude Mythos 5.1. An unnamed White House official told Politico this had been the policy for every new frontier model. Anthropic's model announcement said it was coordinating with the US government to widen access to domestic and international partners. The report says the UK institute's director, in a letter to a parliamentary committee, acknowledged the lack of access to Anthropic's model, said the institute had tested OpenAI's GPT-6 Astra before release, and wrote that it still has pre-release access to some of the most capable models. (Documented fact that The Next Web reports it, moderate: a relay of other outlets' reporting with one unnamed official quoted; no government text was read.) The open threads below list it. The cause of the request is unconfirmed.
Does not prove. That a short window changed any finding. None of the reports read says so.
3. Publication: the evaluated party sees the text first
The second input is the right to publish. Here METR's published terms for two labs can be set side by side.
METR and OpenAI. For GPT-5 (August 2025) and GPT-5.6 Sol (June 2026), METR states that OpenAI's communications and legal team required review and approval of its posts. METR describes an informal understanding that the review checked confidentiality and intellectual property, not risk conclusions. It adds that OpenAI would have had the legal right to block conclusions that depended on non-public information, and that the GPT-5.6 evaluation should not be read as robust formal oversight the public can rely on; the GPT-5 report carries the same three statements. For both reports METR states that it changed no conclusions, takeaways or tone after OpenAI's review. (Documented fact and official claim, high: both METR posts read directly.)
METR and Anthropic. For Claude Opus 5.5 (22 September 2026), METR states that the evaluation ran under an unpaid agreement. METR drafted a summary, Anthropic could review and edit it, and METR signed off on the final text, which is the version in Anthropic's system card. A transparency clause lets METR disclose, without Anthropic's consent, whether Anthropic used its redaction rights and whether findings depended on redacted material. The summary does not say whether Anthropic used its review or edit rights, and it carries no statement matching the no-changes line in the OpenAI reports. (Documented fact, high: METR's post and the system card read directly.)
The two sets of posts tell a reader different things. For OpenAI, METR names the legal right to block and reports no change. For Anthropic, METR may say whether redaction rights were used, and on that question this summary is silent. Whether either lab's agreement carries a term the posts do not describe is unknown. (Inference, moderate.)
System cards speak for the evaluators. OpenAI's GPT-5.6 system card (9 July 2026) has named sections for the UK institute, Apollo Research, Irregular and METR. For METR it says the text is OpenAI's summary of METR's report and links the full report. Anthropic's Opus 5.5 card names the US government's CAISI and METR, and names two contracted red-teaming firms with approximate hours (Trajectory Labs, about 95 hours; 10a Labs, about 56). Neither card states what any evaluator was paid. (Documented fact, high: both PDFs read directly.) Most readers meet the evaluator's work in the evaluated company's words, with the evaluator's own text one link away.
A benchmark example. Epoch AI said on 23 January 2025 that OpenAI commissioned and owns the 300 core problems of its FrontierMath benchmark, has access to problems and solutions apart from a holdout set, and that Epoch needed OpenAI's permission to disclose the sponsorship. Epoch says it received that permission ahead of OpenAI's o3 announcement and then disclosed the partnership, that it did not clarify OpenAI's data access until this post, that many contributing mathematicians did not know of the sponsorship, and that its communication "should have been more systematic." (Official claim, high: Epoch's post read directly.) On Epoch's own sequence, the public read o3's headline result before it knew that the sponsor could see the problems and most of the solutions. (Inference from Epoch's account, moderate.)
Does not prove. That any lab edited, blocked or delayed any finding, or that OpenAI trained on FrontierMath problems or that any score was inflated. Epoch's statement makes no such finding. The records show who holds the right to see and approve text, and in one case that a sponsor's access terms were disclosed after the result.
4. Money: how the checkers are paid
Evaluators answer the money question in different ways, and all four described here publish their answer. That makes this the best documented part of the record.
Grant-funded, with a rule against lab cash. METR says it has not taken funding from frontier AI companies and does not accept donations made by or at the direction of their staff. The same post says frontier companies give it a significant amount of free tokens for evaluations, research and engineering. (Official claim, high: METR funding update of 14 August 2026.) One data point shows the scale of that in-kind support: METR reported that an attacker used a stolen key for three weeks and consumed credits METR values at about $600,000, which the unnamed model developer had granted free. (Official claim, high that METR states it; the dollar figure is METR's estimate.) METR's filed revenue for 2024 was $13,639,155. (Documented fact, high: IRS data through the ProPublica API.) Its About page lists the UK AI Security Institute among its funders. (Official claim, high.) One government evaluator thereby funds a non-government evaluator of the same models.
Grant-funded, with lab employees' money disclosed. Transluce states that in fiscal 2025, 32 percent of its revenue came from the personal holdings of Anthropic employees and 6 percent from those of OpenAI employees, all as unrestricted donations. It says it had done no paid evaluations, may accept compute credits and will disclose them, and recuses staff with a significant financial interest in a system provider. (Official claim, high.) METR's rule forbids that funding pattern. Transluce's rule permits it with disclosure. Transluce's statement is the most specific funding disclosure found in this record.
Paid at market rates. Apollo Research's conflict-of-interest norms say it generally asks the parties it works for to pay fair market value, refuses pay contingent on outcomes, refuses grants or investments from organizations it evaluates or expects to evaluate, and recuses staff with a financial interest in an evaluated organization. Apollo became a public benefit corporation in January 2026 and also builds monitoring products for companies developing or using AI agents. (Official claim and documented fact, high.)
Paid at cost, with a ceiling as an aim. SecureBio says its paid evaluation and audit work is billed at cost, never contingent on results, and that it aims to keep revenue from AI-company services at no more than 25 percent a year. It accepts free API credits. (Official claim, high; the 25 percent is stated as an aim, not a cap.)
Two business models, two ties. Grant-funded nonprofits depend on donors. Fee-based evaluators depend on labs as customers. (Inference, moderate.) Neither model is shown here to produce better or worse evaluations.
The rule no evaluator has. None of the rules read here (METR, Apollo, SecureBio, Transluce, and the AI Evaluator Forum letter in section 9) bars money from a lab's investors, as distinct from the lab or its staff. (Inference from the texts read, moderate.) The overlap is visible in public records. Anthropic's 2021 Series A announcement says the round was led by Jaan Tallinn and names Dustin Moskovitz, Eric Schmidt and the Center for Emerging Risk Research among participants. (Documented fact, high.) The Center for Emerging Risk Research now operates as Macroscopic Ventures, which joined Apollo's 2026 seed round. (Documented fact through secondary sources, moderate.) Schmidt Sciences is a named METR donor, and Schmidt Sciences and Tallinn are philanthropic partners of the industry AI Safety Fund described below. (Official claims, high.) Whether these investors still hold Anthropic equity is unknown.
Lab money reaching safety and policy work. The Frontier Model Forum is an industry 501(c)(6) funded by member fees. Its AI Safety Fund started in October 2023 at more than $10 million, with Anthropic, Google, Microsoft and OpenAI as founding members and four philanthropic partners. An outside administrator ran it until that administrator announced its closure in June 2025; since then the Forum manages it directly. (Official claim and documented fact, high.) In July 2024 Anthropic announced funding for evaluations built by outside organizations, including assessments of its own safety levels, with no amounts or independence terms disclosed. (Official claim, high.) Anthropic and AWS are among the funders of the UK institute's £15 million Alignment Project. (Documented fact as reported, moderate.) One trade report says the OpenAI Foundation's resilience program funds model safety evaluations. (Low; no OpenAI text was readable and no grantee was named.)
Does not prove. That free compute, donations, fees or investor money moved any finding, or that any donor expected anything in return. In-kind support has a value but no documented condition attached to it.
5. Who names its evaluators
A conflict check needs a name to check. OpenAI's GPT-5.6 card and Anthropic's Opus 5.5 card name their outside evaluators (section 3). Google DeepMind's Gemini 3 Pro Frontier Safety Framework report (November 2025) describes chemical, biological, radiological and nuclear testing by unnamed third-party evaluators and names no government institute. xAI's Grok 4.7 model card (21 September 2026) says third-party evaluators with an unrestricted configuration corroborated its internal cyber results, and names none. (Documented fact, high: all four documents read directly.) Later Gemini Pro reports were not reviewed, and Meta's and Microsoft's evaluator disclosures were not reviewed.
Does not prove. That unnamed testing is weaker. Naming is a precondition for an outside conflict check. It does not measure the quality of a test.
6. Governments: tester, promoter and buyer
The United States. The US AI Safety Institute signed agreements with OpenAI and Anthropic on 29 August 2024 for access to major models before and after release; NIST's release gives no financial terms. (Documented fact, high.) On 1 October 2026 the NIST page for the institute's successor carried the title "Center for Advancing Innovation and Standards for Super Intelligence (CAISSI)". It describes the center as industry's primary point of contact in government, and lists evaluations alongside representing US interests against what it calls burdensome foreign regulation. The four 2026 assessments the page lists cover models developed in China. (Documented fact, high: page read directly. The date and instrument of the rename are unknown.) US-model results appear in the labs' own system cards.
One body that both promotes an industry and tests it carries a dual mandate of the kind Congress split in 1974, when it abolished the Atomic Energy Commission. The Nuclear Regulatory Commission's history page says supporters and critics of nuclear power agreed the promotional and regulatory duties belonged in different agencies. (Documented fact, high.) Applying that history to the AI center is this piece's inference. (Moderate.)
The United Kingdom. The UK institute publishes its access requirements, funds METR and runs a fund co-funded by Anthropic and AWS (sections 2 and 4). The same dual-mandate test was not run on the UK side: whether the institute's parent department carries a growth or promotion remit was not checked. The United States reads worse here partly because only its page was read for mandate.
A private evaluator, partly owned by a lab. On 10 February 2025 Scale AI announced an agreement with the US institute to develop evaluations jointly, under which model builders could test with Scale and choose whether to share results with the institute. (Documented fact, moderate.) In June 2025 Meta took a minority stake in Scale, reported as 49 percent and non-voting, for $14.3 billion. (Stake: documented fact, high, from Scale's post. Size: press, moderate.) Whether Scale's evaluator role survived the Meta deal or the rename is unknown.
The government as buyer. Governments also buy models, and a buyer can ask for a different trained behavior. In documents The Intercept obtained through a public-records lawsuit, the Pentagon asked that OpenAI's model "minimize the extent to which" it refuses military requests. Both parties called the documents draft materials provided in error and said the final contract has no such provision. (Documented fact that the drafts say so, moderate; the final-contract statement is the parties' claim, and the final contract was not read.) The same article covers contracts with all four large labs; OpenAI's is the only deployment contract it holds because the government had produced only that one.
Anthropic's own declaration in its court dispute with the Department of War records the same pattern. As the D.C. Circuit's opinion of 25 September 2026 recounts it, commercial Claude refused some national-security tasks in 2024, Anthropic built a government model to perform them, and it came to permit weapons-system design, foreign-intelligence analysis and offensive cyber operations, keeping two exceptions. (Documented fact as the court's account of Anthropic's declaration, high.) The Department then demanded an "all lawful uses" term, Anthropic refused, and the Department excluded it under 41 U.S.C. 4713. The court upheld that exclusion 2 to 1. A parallel action under 10 U.S.C. 3252 was set aside by a federal district court on 27 August 2026, and the appeals court said it had no quarrel with that court's finding that Anthropic showed no bad motive. (Documented fact, high: both courts' records read directly.) In dissent, Judge Henderson argued that under the majority's reading, any supplier told to permit functions the Department deems necessary must agree or risk the same designation. (Documented fact that the dissent argues it, high.) That is a structural point about the leverage a buyer holds over a model's trained limits.
Does not prove. That CAISI tests US models less rigorously, since publication can follow classification and agreement terms; that any government finding tracks a funding tie; that either lab's change to its model was wrong; or that any supplier has changed its terms because of the ruling.
7. People who move between the jobs
Movement is how expertise reaches small evaluators and new agencies. It also carries equity, friendships and job prospects, and the records show it running in every direction.
- Government to lab. Elizabeth Kelly, founding director of the US AI Safety Institute, under whom the 2024 lab agreements were signed, now leads beneficial deployments at Anthropic. (Documented fact, high for the two roles; moderate for the reported dates, early February and mid-March 2025.) The pressure a revolving door creates sits in the earlier role, through the prospect of later work, so the nature of the new job does not settle it. (Inference, moderate.)
- Lab, evaluator, lab governance, government. Paul Christiano led alignment work at OpenAI, founded the Alignment Research Center (whose evaluation team became METR), served as an initial trustee of Anthropic's Long-Term Benefit Trust, and in April 2024 became head of AI safety at the US institute. Anthropic's trust page says he stepped down from the trust that month to take the role. (Documented fact, high for the trust and institute roles; moderate for the earlier background.)
- Funder and lab board to lab. Holden Karnofsky, former co-CEO of Open Philanthropy, which made grants to the Alignment Research Center, served on OpenAI's board from 2017 to 2021 and joined Anthropic in January 2025 to work on its Responsible Scaling Policy. He is married to Anthropic's president, Daniela Amodei. That fact is recorded here because ethics rules commonly treat a spouse's financial interest as one's own. (Documented fact through a secondary aggregator, moderate.)
- Labs to government. Geoffrey Irving, chief scientist at the UK institute, previously led alignment teams at DeepMind and OpenAI. (Official claim from his published bio, high.) Jade Leung, the institute's chief technology officer and then the Prime Minister's AI adviser, was announced on 23 September 2026 as stepping back from both roles at the end of September to become the institute's vice-chair. (Documented fact, high.) A trade outlet calls her a former OpenAI executive. (Moderate.)
- Lab board and evaluator founder. Zico Kolter serves on OpenAI's board, chairs its Safety and Security Committee, and co-founded Gray Swan AI, which built a prompt-injection benchmark with the UK and US institutes; Anthropic's Opus 5.5 card reports results on it. (Documented fact, high for the roles.)
- Lab to evaluator board. Daniel Kokotajlo, who left OpenAI in 2024 and declined its exit non-disparagement terms, became the first holder of a board "mission seat" at Apollo Research in January 2026. (Documented fact, high.) Movement out of a lab can carry criticism as easily as loyalty.
The one published mitigation. The UK government's published declaration for Ian Hogarth, chair of the UK institute, says he agreed on appointment in June 2023 not to hold investments tied directly to its work. He divested holdings including Anthropic at a price fixed at his appointment date, so he gained nothing from later events, and he recuses from decisions at his venture firm on companies within the institute's scope. (Documented fact, high.) It is the most complete conflict record found in this research.
It does not make a fair contrast with the US cases. Hogarth held shares while serving. Kelly's case is a move after service, which federal post-employment law restricts whether or not a document is published. US officials file ethics records with the Office of Government Ethics, and this research did not request them. The US side is unread, not empty.
Does not prove. That any person's work in any role favored a former or future employer.
8. The money above the labs
Evaluators are not the only checkers whose terms matter. Legislators write the rules that decide what labs must disclose, and the labs' investors and political networks fund the people who write them. The records here come from the parties' own filings: SEC reports, federal lobbying disclosures and FEC data. The appendix carries the full figures.
Investors on both sides of the race. Microsoft, Amazon and NVIDIA have each put money, or committed it, into both OpenAI and Anthropic. (Documented fact, high for Microsoft and Amazon; moderate for NVIDIA's OpenAI side.) Microsoft's October 2025 agreement values its OpenAI stake at about $135 billion, roughly 27 percent, and requires an independent expert panel to verify any declaration that OpenAI has reached AGI; in November 2025 Microsoft and NVIDIA committed up to $5 billion and up to $10 billion to Anthropic. (Documented fact and official claim, high.) NVIDIA's OpenAI tie in primary records is a letter of intent to invest up to $100 billion and lease guarantees capped at $105 billion; a reported equity stake rests on press only. (High for the records.) Amazon's June 2026 quarterly report carries Anthropic notes and nonvoting preferred stock at roughly $190 billion combined, and records $15 billion of OpenAI preferred stock bought in early 2026 plus an agreement to buy $35 billion more. (Documented fact, high: filing read directly.) Alphabet's report carries $124.3 billion of non-marketable equity, which it says consists primarily of its investment in one unnamed private company; that the company is Anthropic is an inference, moderate.
A holder of every lab gains whichever lab leads. (Inference, moderate.) The same structure can cut the other way: an investor in every lab has little reason to favor any one lab's account of an incident.
Control rights: a reported early example. The Atlantic, in an excerpt from Kevin Roose's book published on 29 September 2026, reports that during the 2019 Microsoft investment talks Dario Amodei, then at OpenAI, sought to bar Microsoft from vetoing an acquisition of OpenAI, the likely route for the charter's merge-and-assist pledge. It reports Amodei's belief that Sam Altman conceded the term or never asked for it, then lied to him about it. The New Yorker (April 2026) reports that a provision letting Microsoft block a merger was added in June 2019. (Contested account, moderate that such a term existed; low on what was said.) Neither article carries a response from Microsoft or OpenAI on this point; the New Yorker reports that Altman does not remember the exchange. The sources in both sit mostly on one side of the 2020 split that produced Anthropic. Both outlets carry ties of their own: the Atlantic page sells the book through a commission link and says so, and the excerpt does not disclose the Atlantic's 2024 content partnership with OpenAI, whose current status is unknown. (Documented fact, high for what the page shows.) That partnership is with the company the excerpt treats less favorably, so it does not explain the account's direction.
The matching question for Anthropic has no public answer. Whether Amazon's or Google's agreements give either company merger or veto rights over Anthropic is not in any record found; Amazon's filing records part of its stake as nonvoting preferred stock. (Documented fact, high.) The 2025 Microsoft terms, which are public, put an independent panel on the AGI question.
Lobbying, rising from a small base. Anthropic's in-house federal lobbying rose from $0.72 million in 2024 to $3.13 million in 2025, and reached $3.53 million in the first half of 2026 alone. OpenAI's rose from $1.76 million to $2.99 million. Meta ($26.29 million), Amazon ($17.78 million) and Google ($13.10 million) spent more in 2025 on far wider agendas. (Documented fact, high: lda.gov data.) Five companies, Anthropic among them, filed 2025 issue lines that name federal preemption of state AI laws or a moratorium on them; only one line, Google's, states a direction. (Documented fact for the lines, high; grouping them is this piece's reading, moderate.) A lobbying report lists subjects, not positions or outcomes.
Election money, both networks under the same tests. New York's 12th congressional district shows where the two largest AI political networks meet. Alex Bores, the sponsor of New York's RAISE Act, an AI safety disclosure bill, ran there.
- Leading the Future. This super PAC received $75.1 million through 30 June 2026, including $50 million from a16z Capital Management and $25 million from OpenAI president Greg Brockman and his spouse. It made no independent expenditures of its own. It gave $20 million each to Think Big and American Mission, and Think Big spent $8,166,098 opposing Bores through 21 August 2026. (Documented fact, high: FEC data.) Stated reason, from its site: build lasting political infrastructure for AI by backing pro-AI candidates. (Official claim.)
- Public First. Anthropic gave Public First Action, a 501(c)(4), $40 million in 2026 and says the money cannot be used to influence any candidate's election. (Official claim, high.) FEC records show Public First Action moved at least $24.25 million into three super PACs between December 2025 and July 2026, and one of them, Jobs and Democracy PAC, spent $15,026,938 supporting Bores. (Documented fact, high.) The records do not show whose money Public First Action moved, and a restricted gift and transfers from other money can both be true. Stated reason, from Anthropic's posts: frontier transparency rules, and no preemption of state laws unless Congress enacts stronger safeguards. (Official claim.)
- Both. Donors listing Anthropic as employer, its chief executive among them, gave $3,154,900 of the Public First super PAC's $4,909,900 in itemized receipts. Employees of at least four labs appear as donors across the two networks, which also backed the same Oklahoma candidate. Think Big, American Mission and Jobs and Democracy PAC each paid nearly all of its independent spending to one vendor; Defending Our Values PAC used 15 payees. A Campaign Legal Center complaint alleges that the vendors used by Think Big and its sister committee hide the true payees; no answer or FEC action was found, and no complaint was found about any other vendor. (Money: documented fact, high. Complaint: contested account.) The complaint's figure of more than $125 million traces to its own donor list, which counts the a16z partners' memo attributions as extra money; the FEC record supports $75.1 million.
- Corporate PACs. Anthropic's own PAC, registered in April 2026, gave $156,000 to committees of both parties by 31 August. No committee is named for OpenAI or xAI. (Documented fact, high.)
Does not prove. That any stake, lobbying contact or contribution changed a safety, disclosure or evaluation decision, or any vote. No record read shows an Anthropic dollar reaching a super PAC. The stated reasons are the parties' words.
9. Older checkers, and what the evaluators now ask for
Checkers paid by the checked, before AI. Auditors and credit raters have long been paid by the companies they review, and their records show what that dependence can produce.
- The reviewer the client wanted gone. In 2002 congressional investigators released Arthur Andersen notes showing that Carl Bass, Andersen's own technical reviewer on the Enron account, objected to Enron's accounting, and that a handwritten note of 12 March 2001 reads "Client sees need to replace Carl." He was removed from the Enron engagement. (Documented fact as reported by the Associated Press, moderate to high.)
- The rater paid by the issuer. Under the issuer-pays model, the company whose debt is rated pays the rating agency. In January 2017 Moody's agreed to pay $863,791,823 to settle with the Justice Department and state partners, accepted a Statement of Facts, and took on five years of compliance commitments. Moody's said the settlement contained no finding or admission of a legal violation. (Documented fact, high: settlement agreement read directly.)
- The checker who knew. The SEC's inspector general found that the agency's Fort Worth examiners concluded in 1997, 1998, 2002 and 2004 that Stanford Financial's certificates of deposit were likely a Ponzi or similar scheme, and that no meaningful enforcement effort followed until late 2005. The office's former head of enforcement sought three times to represent Stanford after leaving and did so briefly in 2006. (Documented fact, high: inspector general report OIG-526, read directly.)
Where the comparison breaks. Auditors and raters work under statutes, licensing and liability that AI evaluators lack. That gap is the point of the comparison: no statute read for this piece sets duties for AI evaluators of the kind those professions carry. (Inference, moderate.)
Does not prove. That Andersen's removal of Bass caused Enron's collapse, that issuer-pays shaded any named rating, or any motive at the SEC. Nothing here shows that any AI evaluator has behaved as these checkers did.
September 2026: the labs offer embedded evaluators, and the evaluators set conditions. Dario Amodei's essay "We Must Pace the Frontier", dated September 2026, commits Anthropic to give a team of third-party evaluators ongoing, employee-like access and the right to publish key findings without Anthropic's editorial control. Redaction is limited to four categories (security-sensitive, legally privileged, commercially sensitive and third-party confidential material), a power the essay itself calls narrow, and the essay says findings cannot be redacted for being unfavorable. (Official claim, high: essay read directly.) Anthropic's 18 September post names Faculty, Accenture's AI business, as the first embedded partner, says Anthropic will fund it directly and that each company expects to invest at least $1 billion in the area over five years, and states no publication terms. (Official claim, high.) TechCrunch reported that Sam Altman said OpenAI would also embed evaluators, and that neither company had named evaluators, timing, scope or disclosure rules. (Moderate; OpenAI's own statement was not read.)
On 18 September the AI Evaluator Forum, whose members include METR, RAND, SecureBio and Transluce, hosted an open letter signed by 200 or more people in a personal capacity. It sets minimum conditions: embedded evaluators should not be owned or governed by frontier companies, should have no other significant commercial business with them, should take no pay contingent on findings, and need privileged-employee access and protection from retaliation. (Official claim, high for the letter text.)
Does not prove. That either lab's commitment will meet the Forum's conditions. No embedded-evaluation contract is public, so how the redaction power will be used is unknown. The Forum's commercial-business clause would bear on paid evaluators such as Apollo, and its ownership clause on an evaluator like Scale.
10. The test: published terms anyone can audit
The record above does not show a corrupted evaluation. It shows a set of dependencies that are mostly invisible from outside. The same small set of facts settles most of the questions this piece raised, and for most engagements no one publishes them. (Inference, moderate.) A reader auditing any check of a frontier model needs to know:
- Who paid, and how much, in cash, credits and staff time, including money from the lab's staff and investors.
- What access was granted, for how many days, and which versions of the model.
- Who could review or edit the text, under what rights, and whether those rights were used.
- Who owns or funds the evaluator, and whether it sells anything else to the lab it checks.
- Who on the team recently worked for, or holds a stake in, the lab, and what recusal followed.
Measured against that list, the record is uneven. METR publishes its funding rule and, for the three engagements read here, the review terms. Transluce publishes the share of its revenue from each lab's employees. Apollo and SecureBio publish conflict rules. The UK government published its chair's divestment. OpenAI and Anthropic name their evaluators. Google DeepMind and xAI, in the documents read, do not. No lab document read states what the lab pays any evaluator. No embedded-evaluator contract is public. The US ethics records for the people who moved were not requested, and so remain unread. (Each item documented above; the scoring is inference, moderate.)
The test asks for disclosure. Every evaluator with access depends on the lab that grants it, so the question for each engagement is whether the terms of that dependence can be read.
Does not prove. That evaluators with fewer ties produce better results, or that publishing terms would change any finding. The list is a proposal; no one has tested whether it improves outcomes.
How this connects to Who Knew First
Who Knew First is a record of nine 2026 incidents in which an AI agent crossed a boundary. It found that each review with access to an operator's logs was agreed with the party under review, and it scored that dependence one engagement at a time. This piece widens the frame from a single incident review to the economy of evaluation as a whole: who pays, who grants access, who approves the text and who moves between the jobs. It leaves that piece's conclusion as it stands. A dependence is a reason to publish the terms, and nothing here shows that any finding was shaded. A reader who has not read Who Knew First needs only this: whoever holds the logs and the money also holds most of the inputs a checker works from, so the terms under which a checker gets them are part of the evidence.
What this does not prove
- That any evaluation, by any party named here, was softened, edited, delayed or shaded.
- That any stake, grant, credit, fee, lobbying contact or political contribution changed any decision by a lab, an evaluator, an agency or a legislator.
- That evaluators with fewer ties produce better results, or that any business model for evaluation is superior.
- That any person named acted against the public interest, or anything about any person's or company's motives. The stated reasons quoted are the parties' own words.
- How common undisclosed ties are. The research was targeted, so a missing row means "not found", not "absent".
- That the reported accounts (the Atlantic and New Yorker items, the UK access report, TechCrunch's windows) are accurate. They are labeled as reported, with the responses found.
- That the confirmed records are true. A confirmed claim matches its source; it does not prove the source right.
Open threads: where others can dig
Each item below was found and not settled. It is listed with what is known, what limits it, and the record that would settle it. Anyone can pick one up.
- SEC orders on false AI claims. The SEC settled with Delphia and Global Predictions (2024) and Presto Automation (2025) over misleading statements about AI, each without admission or denial. These were read only through summaries because sec.gov refused a direct fetch. Read the orders directly. No regulator finding of a false capability claim against a frontier lab was found in the SEC and DOJ sites; the FTC was not searched.
- Irregular's investors and customers. Irregular, an evaluator formerly called Pattern Labs, raised $80 million in a round led by Sequoia Capital and Redpoint Ventures and says it works with OpenAI, Anthropic and Google DeepMind. TechCrunch, citing the Financial Times, reported that Sequoia holds stakes in OpenAI and xAI and was joining an Anthropic round (moderate). Whether the labs pay Irregular, and how much, is unconfirmed.
- Why the UK institute lost access. The reported White House request that labs withhold models from the UK institute rests on a relay of Politico and Bloomberg (moderate). Find a government text, the director's letter to the parliamentary committee, or a lab statement on the record.
- METR's critics, held to the same standard. A September 2026 report collected criticism that METR is too close to Anthropic's investors and staff, including from David Sacks and Perry Metzger. NPR reports that Sacks served under two ethics waivers; that report and its divestment figures reached this research only through a search summary (moderate). Read the Office of Government Ethics letter before using any figure, and find who funds Metzger's organization.
- More people in the movement map. Paul Nakasone joined OpenAI's board and safety committee in 2024; Anthropic formed a national-security advisory council of former officials in 2025; Michael Kratsios moved from Scale AI to the White House; one trade report says OpenAI co-founder Wojciech Zaremba heads the OpenAI Foundation's resilience grants (low). These were not re-read for this piece (moderate, Zaremba low). Request the Office of Government Ethics records for Kelly, Christiano and Kratsios.
- A contested pass-through. A September 2026 post says METR's funding includes $10 million passed through a RAND program funded by Coefficient Giving; a commenter disputes it (low). Only RAND's or METR's own records can settle it.
- Press-only figures. NVIDIA's reported $30 billion OpenAI equity stake, Google's reported plan to invest up to $40 billion in Anthropic, and Meta's reported $45 million and $65 million for its state-level political committees all rest on press reports. Find the filings.
- The 2017 funding plan. The Atlantic (29 September 2026) and The New Yorker (April 2026) both report that OpenAI executives pitched raising money in 2017 by offering governments rights to future AGI; the Atlantic names the United States, China and Russia, and the New Yorker names China and Russia. OpenAI told the Atlantic it briefly weighed an international cooperative and "did not consider selling AGI to China or Russia"; the New Yorker reports Brockman's position that he never seriously entertained an auction. The two accounts differ on who devised the plan and how it ended, and their sources likely overlap and sit mostly on one side of the later split (contested account; moderate that a plan was pitched, low on its terms). The matching question for Anthropic's own fundraising, including any offers to foreign investors, has not been researched here. Ask the same question of every lab's fundraising with the same verbs.
- Records to pull. IRS e-file data for METR, SecureBio, the Alignment Research Center, Redwood Research, the Frontier Model Forum and Apollo's foundation; Kevin Bass's public
metr-deeprecords pack (read its claim gate before reusing a figure); the AEF-1 standard clause by clause against each evaluator's policy; CAISI's agreement texts by public-records request; on-the-record requests to Google DeepMind and xAI to name their evaluators; and the embedded-evaluator contracts once signed, tested against the Forum's conditions. - Two asymmetries this research could not close. OpenAI's website refused every read, so OpenAI's own stated reasons are missing from three investor rows, while Anthropic's own descriptions appear in several. And the UK institute's parent department was not checked for a promotion remit, so the dual-mandate test ran on the US side only.
How this was made
An Anthropic-built model, Claude Opus 5.5, drafted this piece and re-checked its sources at the author's request. That is not an outside review. Anthropic is a party in many of the rows: as an evaluated lab, an investor-backed company, a funder of political advocacy and a funder of evaluation work.
The research behind it read primary records where they could be reached: system cards, evaluator posts and policies, SEC filings, federal lobbying data, FEC data, court opinions, government pages and an inspector general's report. A second pass re-read about 375 claims in the two underlying ledgers and found 14 problems, all corrected before publication; most sat on rows first read through a summarizing tool. A fairness pass looked for unequal standards between labs and states and corrected what it found. A third pass applied a swap test to the Anthropic items in the research files and in the body of Who Knew First, replacing one lab's name with another's and asking whether the wording would change; it found nine places where the wording or the selection of facts read more favorably to Anthropic. The research files were corrected. On 1 October 2026, a dated follow-up on Who Knew First published five corrections to that page; four of them came from this pass. About 900 Anthropic lines in that page's ledgers have not been read in full by any pass. All three passes were run by the same model family that compiled the record, on samples chosen by risk. A same-maker check can miss what an outside reader would catch. The items naming Anthropic are the ones most in need of an outside check, and corrections from any reader are welcome.
Two further conflict channels belong here. Anthropic funds Public First Action, whose network spent $15.03 million supporting the sponsor of New York's AI disclosure bill (section 8), so a finding that favors disclosure rules favors a policy the drafting model's maker publicly backs. And the author's own verification tooling is aimed at frontier-lab evaluation work, so the embedded-evaluator commitments described in section 9 would create demand for it. The author has also sent review material to METR, one of the evaluators examined here. None of these channels shows that the record leans. Each is a reason for the outside check this note invites.
No source was reached by getting around a paywall, a sign-in or a bot check. Pages that refused access are marked unread. Quotations from copyrighted sources are limited to one short phrase per source.
Corrections
None yet. Corrections will be listed here with their dates.
Appendix: the ledger, one row format for every party
Every party gets the same columns. Anthropic is listed first because the drafting model's maker is the party most in need of an outside check. Each cell repeats or extends a claim from the sections above. The Sources section lists the records behind every cell. "Not found" means not found in this research, not absent.
Labs
| Party | Evaluators, access and publication | Money to checkers | Investors and control | Political money and lobbying | People | Not found |
|---|---|---|---|---|---|---|
| Anthropic | Names CAISI, METR and two contracted red-team firms in its Opus 5.5 card; may review and edit METR's summary under a clause that lets METR say whether redaction rights were used; the Opus 5.5 summary is silent on use. Built a government model that does tasks commercial Claude refused [court record] | Employees' personal holdings: 32% of Transluce's FY2025 revenue; funds third-party eval development, amounts undisclosed; co-funds the UK Alignment Project; AI Safety Fund founding member; funds Faculty, with at least $1B each from Anthropic and Accenture expected over five years | Amazon notes and nonvoting preferred about $190B on Amazon's books; Microsoft up to $5B, NVIDIA up to $10B; Google's large unnamed stake inferred | $40M to Public First Action, whose network moved at least $24.25M into super PACs; employees and CEO gave $3.15M to Public First; AnthroPAC $156K to both parties; lobbying $3.13M (2025), $3.53M (H1 2026) | Hired US institute's founding director; a trust member left for the US institute; Karnofsky joined | Pay to contracted testers; merger or veto rights held by its investors |
| OpenAI | Comms and legal approve METR's posts, with a stated legal right to block conclusions resting on private information; METR reports no change to conclusions; card summarizes METR in OpenAI's words. Pentagon draft asked its model to refuse less; parties say final contract omits it [The Intercept] | Employees' personal holdings: 6% of Transluce's FY2025 revenue; Foundation reportedly funds safety evaluations (low); AI Safety Fund founding member | Microsoft about 27% with an AGI verification panel; NVIDIA letter of intent and $105B lease guarantees; AMD warrant; Amazon $15B preferred bought and $35B more agreed. Reported 2019 Microsoft merger-block term (contested) [Atlantic, New Yorker] | President and spouse gave $25M to Leading the Future, whose affiliate spent $8.17M against the RAISE Act's sponsor; no PAC under its name; lobbying $2.99M (2025) | Board member co-founded Gray Swan; former staff at UK institute | Pay to any evaluator; Foundation grantee list; its own stated reasons (site unreadable) |
| Google and Google DeepMind | Gemini 3 Pro safety report: unnamed third-party evaluators, no institute named | AI Safety Fund founding member | Large unnamed private stake and $20B milestone commitment; press ties it to Anthropic | Google NetPAC $1.51M; lobbying $13.10M (2025), the one preemption line stating a direction | Former DeepMind lead at UK institute | Evaluator names and terms |
| xAI | Grok 4.7 card: unnamed third-party evaluators | None found | Not gathered | No LDA filings or FEC committee under its name | None recorded | Evaluator names; investor terms |
| Meta | Not reviewed | None found | Minority stake in Scale AI (press: 49%, $14.3B) | Meta Platforms PAC $0.30M; lobbying $26.29M (2025); state committees, amounts press-only | Scale's founder joined Meta | Evaluator disclosures |
| Microsoft | Not reviewed | AI Safety Fund founding member | Stakes in OpenAI and commitment to Anthropic | MSVPAC $1.24M; lobbying $9.36M (2025, different basis) | None recorded | Evaluator disclosures |
Evaluators
| Party | Funding rule | Lab money or credits | Publication and access terms | People and ties | Not found |
|---|---|---|---|---|---|
| METR | No company money, no staff donations | Free tokens; one stolen key used about $600K of donated credits; donors include the UK institute and Schmidt Sciences (see section 4 on investor overlap) | Review rights for OpenAI (approval) and Anthropic (edit), with different disclosure terms | Board and advisors listed | Per-donor amounts; total in-kind value; 2024 Form 990 text |
| Transluce | Accepts lab-employee donations, disclosed | 32% Anthropic and 6% OpenAI employees' holdings (FY2025) | No paid evaluations to date; will disclose credits | Chairs the AI Evaluator Forum | Value of any credits used |
| Apollo Research | Market-rate fees; no outcome-contingent pay; no grants or investment from evaluated labs | Fees from labs (amounts unknown) | Not stated | Seed investor Macroscopic Ventures is the renamed Center for Emerging Risk Research, an Anthropic Series A participant (moderate); former OpenAI staffer holds mission seat | Revenue share from labs |
| SecureBio | At-cost services; AI-company revenue at most 25% a year as an aim | Free API credits | Not contingent on results | Recusal for recent lab work | 2024 and 2025 shares |
| Irregular | Not stated | Says it works with three labs; payment unconfirmed | Not stated | Lead investor holds OpenAI and xAI stakes (moderate) | Whether labs pay it |
| Scale AI | Not stated | Meta minority stake | Joint evaluation agreement with the US institute (2025) | Former managing director now at the White House | Whether its evaluator role continues |
| Epoch AI | Not reviewed | OpenAI commissioned and owns FrontierMath's core problems | Needed OpenAI's permission to disclose the sponsorship | None recorded | Funding beyond the OpenAI commission |
Governments
| Party | Mandate | Access | Money ties | People | Not found |
|---|---|---|---|---|---|
| US (CAISI, renamed CAISSI) | Industry contact, evaluation, and advocacy against foreign regulation in one body | Voluntary agreements; 2026 listed assessments cover Chinese models | None found | Founding director to Anthropic; Christiano from Anthropic's trust; Scale's managing director to the White House | Agreement texts; US-model findings; ethics records (not requested) |
| UK (AISI) | Parent remit not checked | Published access requirements; reported loss of access to one Anthropic model | Funds METR; runs a fund co-funded by Anthropic and AWS | Chair's published divestment; staff from OpenAI and DeepMind | Payment terms with labs, if any |
Sources
Read directly unless marked. "Summary" means the text reached this research only through a search or summarizing tool and carries moderate weight at most. "Unread" means located but refused or not opened.
History and older checkers
- Dwight D. Eisenhower, Farewell Address, 17 January 1961. Avalon Project, Yale Law School (public domain).
- SEC Office of Inspector General, Investigation of the SEC's Response to Concerns Regarding Robert Allen Stanford's Alleged Ponzi Scheme, Report OIG-526, 31 March 2010. https://www.sec.gov/files/oig-526.pdf
- Moody's FIRREA settlement agreement with the Justice Department and state partners, January 2017 (justice.gov PDF; the press page and annexes unread behind a bot check). CNN, 13 January 2017.
- Associated Press, "Enron silenced a chief critic at Andersen", via Deseret News, 3 April 2002. https://www.deseret.com/2002/4/3/19646977/enron-silenced-a-chief-critic-at-andersen/
- US Nuclear Regulatory Commission, History. https://www.nrc.gov/about-nrc/history.html
Labs
- Anthropic, Claude Opus 5.5 System Card, 22 September 2026.
- OpenAI, GPT-5.6 System Card, 9 July 2026. https://deploymentsafety.openai.com/gpt-5-6/gpt-5-6.pdf
- Google DeepMind, Gemini 3 Pro Frontier Safety Framework Report, November 2025.
- xAI, Grok 4.7 Model Card, 21 September 2026.
- Dario Amodei, "We Must Pace the Frontier", September 2026. https://darioamodei.com/post/we-must-pace-the-frontier
- Anthropic, embedded evaluation with Accenture, 18 September 2026. https://www.anthropic.com/news/accenture-embedded-evaluation
- Anthropic, "A new initiative for developing third-party model evaluations", 1 July 2024.
- Anthropic, "The Long-Term Benefit Trust", 19 September 2023, with later updates.
- Anthropic, Series A announcement, 28 May 2021.
- Anthropic, donations to Public First Action, 12 February and 21 July 2026.
- Anthropic, compute and investment announcements of 22 November 2024, 23 October 2025, 18 November 2025, 6 April 2026, 20 April 2026 and 28 May 2026 (Series H).
- Microsoft, "The next chapter of the Microsoft-OpenAI partnership", 28 October 2025; Form 10-K for the year to 30 June 2026, Note 3.
- NVIDIA, OpenAI partnership release, 22 September 2025; Forms 10-Q for quarters ended 26 April and 26 July 2026.
- AMD, OpenAI partnership release, 6 October 2025.
- Amazon, Form 10-Q for the quarter ended 30 June 2026, Note 2.
- Alphabet, Form 10-Q for the quarter ended 30 June 2026.
- Scale AI, company announcement, 12 June 2025; US AI Safety Institute agreement post, 10 February 2025.
- Epoch AI, "Clarifying the creation and use of the FrontierMath benchmark", 23 January 2025. https://epoch.ai/latest/openai-and-frontiermath
Evaluators
- METR: funding update, 14 August 2026; security update, 31 August 2026; Claude Opus 5.5 summary, 22 September 2026; GPT-5.6 Sol summary, 26 June 2026; GPT-5 report, 7 August 2025; About page. https://metr.org
- Apollo Research: conflict-of-interest norms, 26 November 2025; public benefit corporation announcement, 20 January 2026. https://www.apolloresearch.ai
- Transluce, Independence and Transparency Policy, updated 27 August 2026.
- SecureBio, Conflicts of Interest Policy.
- AI Evaluator Forum, embedded-evaluation letter, 18 September 2026, and AEF-1 page (the AEF-1 PDF unread). https://aievaluatorforum.org
- Frontier Model Forum, About and AI Safety Fund pages.
- IRS data via the ProPublica Nonprofit Explorer API for METR, SecureBio and others.
- Irregular funding release, 17 September 2025 (summary). TechCrunch citing the Financial Times on Sequoia, 18 January 2026 (summary).
- Kevin Bass,
metr-deeppublic-records pack, 16 September 2026 (README read; figures not reused).
Governments
- NIST, US AI Safety Institute agreements with OpenAI and Anthropic, 29 August 2024; NIST CAISI (CAISSI) page, read 1 October 2026; appointment release of 16 April 2024.
- UK AI Security Institute, "Early lessons from evaluating frontier AI systems", 24 October 2024.
- UK Government, Ian Hogarth's declared outside interests, updated 21 February 2025.
- UK Government, Jade Leung appointment, 23 September 2026.
- Anthropic PBC v. Department of War, D.C. Cir. No. 26-1049, opinion of 25 September 2026. https://media.cadc.uscourts.gov/opinions/docs/2026/09/26-1049-2194984.pdf
- Anthropic PBC v. Department of War, N.D. Cal. No. 3:26-cv-01996, orders of 26 March and 27 August 2026.
Money and politics
- FEC data pages and API, processed data for 2025 to 2026: Leading the Future (C00916114), Think Big (C00923417), American Mission (C00916692), Public First (C00930503), Jobs and Democracy PAC (C00928374), Defending Our Values PAC (C00928390), AnthroPAC (C00946111), and the corporate PACs of Amazon, Google, Microsoft, Meta and NVIDIA. Read 1 October 2026.
- Lobbying Disclosure Act filings, lda.gov public API, filing years 2023 to 2026, read 1 October 2026.
- Campaign Legal Center, complaint to the FEC against American Mission and Think Big, May 2026.
- New York Assembly bill A6453 (2025), sponsor record. nysenate.gov
- Leading the Future, Public First Action and Public First websites.
Press and reported accounts
- Sam Biddle, The Intercept, 8 September 2026. https://theintercept.com/2026/09/08/military-ai-weapons-contracts-openai-anthropic-google/
- TechCrunch, "Anthropic and OpenAI want to embed safety evaluators. Will they really be independent?", 16 September 2026.
- The Next Web, report relaying Politico and Bloomberg on UK institute access, September 2026 (Politico unread).
- Kevin Roose, "Inside the Biggest Feud in Artificial Intelligence", The Atlantic, 29 September 2026, adapted from his book The AGI Chronicles. https://www.theatlantic.com/technology/2026/09/openai-v-anthropic-inside-biggest-rivalry-tech/688819/ Read in a browser through the publisher's own gift link.
- Ronan Farrow and Andrew Marantz, profile of Sam Altman, The New Yorker, online 6 April 2026, issue of 13 April 2026 (read by keyword search, not end to end).
- Nieman Lab, Sarah Scire, on The Atlantic's OpenAI partnership, 5 June 2024.
- DefenseScoop, Pentagon AI contracts, 14 July 2025.
- Wikipedia, Holden Karnofsky (secondary aggregator, moderate).
- Summary only: CNBC and TechCrunch on Meta and Scale (June 2025); reports on Kelly's dates, Kratsios's confirmation, Leung's OpenAI role and the May 2026 CAISI agreements; trade press on the OpenAI Foundation; NPR on David Sacks; SEC releases on Delphia, Global Predictions and Presto.
Unread, refused or behind a wall (not bypassed): openai.com (every page), politico.com, ProPublica's Form 990 PDFs, commerce.gov, HPCwire, The Hill, Fast Company, the Senate Banking letter on David Sacks, and Coefficient Giving grant pages.