David Gringras
My official training spans medicine, law, and health policy; my research has wandered across AI evaluation, compute governance, tort liability for AI, and international governance.
My first work in the field was on the institutional questions of the AISI Network: whether mutual evaluation in the style of the FATF could transfer, and how a set of red lines could be enforced without formal treaty-making authority. Through that work, I came across the Network’s first joint testing exercise, in which three different institutes ran the same benchmark on the same model using the same prompting strategy—and got different results. Small differences in methodology, like output parsing and token limits, can shift outcomes. From a biomedical perspective, three different labs running the same assay on the same sample and getting different results is a major methodological crisis; here it was treated as a footnote. I could not find any literature that adequately settled the questions this raised.
So I tested it. In a pre-registered experiment, I demonstrated that format alone could shift measured safety by five to twenty points. Sandbagging and evaluation-awareness are well documented in frontier models; strategic underperformance sits comfortably inside a noise floor that wide. A model that fails an evaluation gives you grounds for alarm; a model that passes does not certify anything. Which is a problem, because nearly every safety-critical trigger in current governance frameworks routes through evaluation results. So I turned to design, mapping out the precise ways in which those frameworks depend on unreliable tests, and arranging them into a taxonomy that practitioners can actually use. I proposed that any mechanism whose failure could run to catastrophe should have at least one activation pathway that owes nothing to evaluation results. And because those activation pathways need to rely on levers outside the model itself, the compute substrate: collecting and cross-referencing data on AI infrastructure that was previously scattered across separate sources.
There is also a separate, medical thread. When a friend asked me about a medical concern (something one gets used to quickly as a physician!) and I wanted to ensure my advice was reasonable since I had not practiced medicine in several months, I presented the case to a frontier LLM. The diagnosis and advice aligned with my own, but was far more detailed and, frankly, overall much better than my own written response would have been. Feeling slightly inadequate (but pleased about no longer needing to be my friends’ personal doctor), I sent a screenshot and suggested the LLM as a reasonable first port-of-call in future. My friend was confused; they had asked the exact same model about the exact same problem with no fewer clinically relevant details and had only approached me because they did not get a helpful answer. My clinically-trained language alone had produced a meaningfully different response. I tested this crudely a couple of times myself, and decided another study was warranted.
I then demonstrated that on identical medical presentations, models withhold safety-critical clinical information from users they read as laypeople while providing it to users they read as clinicians. A harm created by the safety measure itself, and entirely invisible to evaluations built around dangerous outputs. Safety training is scored almost entirely based on commission, what a model says that it should not; omission, what it withholds that it should provide, is nearly free. This was accompanied by legal work on the nature of the tort exposure that omission in fact creates, and the incentive landscape that entrenches the behaviour.
Both chains of work end the same way medicine and law learned to end them. Pathways redesigned around unreliable tests, pharmacovigilance for harms caused by the intervention itself, and legal doctrine for harms of omission. The most reliable path I have found from research to impact is through things people can actually use: frameworks, taxonomies, datasets, legal arguments.
- EnforcementRed lines & the AISI Network
- Deployment gatesOrion AI Governance Initiative
- LiabilityDefensive AI
- Benchmark scoreSafety Under Scaffolding
- Model behaviourIatroBench
- Published evidenceFrontier Lag
- Model internalsno audit of mine
- Training corpusno audit of mine
- Compute substrateScrutica
Evaluation reliability
- Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety
- A pre-registered factorial experiment: six frontier models; four scaffold architectures; four safety benchmarks; >62,000 primary observations (86,000+ in total). Tests whether scores on ‘safety’-type benchmarks are robust to changes in the way they’re assessed. Mostly not – simply switching the measurement format is enough to move a model’s apparent safety score by 5–20 percentage points, matching or exceeding the gap between models themselves. No composite score cleared a point estimate of zero measurement reliability. Cited by agentic-safety work reproducing the scaffold effect and extending it to real-world harm.
- IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
- IatroBench found that frontier models furnish materially superior clinical guidance to users inferred to be physicians than to those perceived as laypersons on identical symptom presentations (p=0.003), which the paper attributes to three mechanistically distinct failure modes (at least two of which current evaluation frameworks would not detect).
- Frontier Lag: A Bibliometric Audit of Capability Misrepresentation in Academic AI Evaluation
- A co-authored bibliometric audit of the applied-domain capability evaluation literature intended to measure how far academic capability claims trail the frontier. We find the median paper evaluates a model around a year behind the frontier (a gap that is widening and only ~25% attributable to peer-review/publication delay), with most abstracts generalising to “AI” rather than the model and configuration actually tested.
Policy & governance
- Strengthening the International Network for Advanced AI Measurement, Evaluation and Science to Support Enforceable Global AI Red Lines
- Evaluations and Collaborations Lead on a team examining the transferability of FATF-style international governance to the AISI Network, as a potential model for red-lines enforcement without treaty authority.
- Defensive AI: When Safety Alignment Creates Tort Liability for Medical Information Omission
- Legal analysis arguing that RLHF-trained safety measures create omission-based tort exposure that is structurally harder to defend than the commission errors the safety training was designed to prevent. The paper traces how liability structures and emerging AI regulation create incentive gradients that entrench this pattern.
- Orion AI Governance Initiative
- Lead a research project I conceived on eval-independent governance that maps safety-critical triggers across frontier governance frameworks, marks which route through model-evaluation results, and catalogues how evaluations fail in practice (eval awareness, developer and model sandbagging, measurement-reliability breakdown) to inform a taxonomy and design principles for eval-independent triggers. Paper in final preparation for submission.
- Scrutica
- I solo-built it because many such eval-independent levers require cross-referenced compute-infrastructure data that was previously scattered across separate sources. It is an open-access compute-governance platform bringing compute facilities, GPU supply chains, export controls and sovereign-AI programmes into one queryable system; the aim simply being to remove that data bottleneck.
Full research & publications →
Background

I believe my (somewhat unusual) background offers a comparative advantage, as it allows me to look at AI risk through the lens of “mature” high-stakes fields, which have a lot to offer the comparatively institutionally immature field of AI safety/governance: medicine trained me to think about high-stakes decisions under uncertainty; law trained me to think about incentives, liability, and institutional failure; and health policy trained me to care about whether theoretical interventions survive contact with real systems.
- MBChB, Medicine & Surgery (Edinburgh) · MA, Law (The University of Law) · MPH, Health Policy (Harvard)
- Frank Knox Memorial Fellowship Harvard University
- Peer-reviewed medical publications British Journal of Cardiology · European Journal of Public Health · Lifestyle Medicine
- First Prize, Tort Law Awards 2024 Thomson Reuters & 12 King’s Bench Walk
- Best Student & Best Dissertation, Global Health Policy University of Edinburgh