In progress

Research

Projects & initiatives

Arcadia Impact AI Governance Taskforce · The Future Society
My first formal experience in the field, as Evaluations and Collaborations Lead on a team examining the transferability of FATF-style international governance to the AISI Network, as a potential model for red-lines enforcement without treaty authority. After six months of work, the paper has recently been finalised and submitted for publication.
Research fellowthwingetal.app2026
Orion AI Governance Initiative
The Safety Under Scaffolding findings, together with increasing documentation of evaluation-awareness/sandbagging in frontier models, paint a bleak picture of the reliability of evaluations and, by extension, evaluation-gated governance. This prompted me to conceive a project mapping where governance frameworks’ safety triggers depend on evaluation results, analysing their reliability, and developing a taxonomy of evaluation-independent governance mechanisms, built as a practical resource for governance practitioners. I have led a competitively selected team of fellows on this project.
Project supervisor2026–
Scrutica
I solo-built it because many such eval-independent levers require cross-referenced compute-infrastructure data that was previously scattered across separate sources. It is an open-access compute-governance platform bringing compute facilities, GPU supply chains, export controls and sovereign-AI programmes into one queryable system; the aim simply being to remove that data bottleneck.
Founderscrutica.com2026–

Papers

Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety
Through the Arcadia research, I came across the Network’s first joint testing exercise, where three institutes ran the same benchmark on the same model with the same prompting strategy and got different results because of small methodological variations. From my biomedical perspective (where measurement consistency is considered a prerequisite), this seemed a bigger deal than it had been treated as, and because I could not find existing literature that settled the questions it raised, I investigated. Through this sole-authored, pre-registered study across six frontier models and 86,000+ observations, I found that model safety scores are significantly contingent on measurement conditions in inconsistent and statistically unpredictable ways (G = 0.000); a scaffold that shifts one model’s safety score can have an entirely different effect on another (since independently replicated, arXiv:2604.01438). The paper provides policy recommendations and a reusable framework for testing evaluation sensitivity.
Pre-printarXiv:2603.100442026
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
My attempt to introduce the medical concept of iatrogenic harm (injury caused by the apparatus that was meant to help) to AI safety. Safety benchmarks almost all score commission harm (e.g. a model giving incorrect/dangerous medical advice) so labs optimised to pass them, hedging and refusing, to look safe and limit liability. But, in certain contexts, withholding information can be just as dangerous as providing incorrect information. IatroBench tested six frontier models in sixty such scenarios (medical queries where there is safe, patient-actionable information to provide, and where not providing this would all-but guarantee harm), finding significant identity-contingent withholding (p=0.003): models systematically withheld life-threatening clinical information from perceived laypersons whilst providing it (in identical scenarios) to physicians, lawyers, or jargon-fluent laypeople (i.e. the trigger is perceived status rather than professional appropriateness).
Pre-printarXiv:2604.077092026
Defensive AI: When Safety Alignment Creates Tort Liability for Medical Information Omission
The legal companion to IatroBench, in which I argue that this omission harm introduces tortious liability. I wrote this in the hopes of drawing attention to the less-obvious legal exposure introduced by this less-obvious vector of harm, which I see as a direct consequence of a misaligned legal incentive landscape where the status-quo is expected-value-maximising; safety training penalises commission and leaves omission almost free because only the former is presumed to introduce liability exposure. Beyond this narrow context, I see a lot of potential for tort law/common law to play a valuable defence-in-depth function in aligning legal incentives with safety (in a way that is more effective at discouraging narrow ‘compliance-on-paper’ specification gaming than rigid statutory law).
Working paperSSRN 64020982026
Frontier Lag: A Bibliometric Audit of Capability Misrepresentation in Academic AI Evaluation
A co-authored bibliometric audit of the applied-domain capability evaluation literature intended to measure how far academic capability claims trail the frontier. We find the median paper evaluates a model around a year behind the frontier (a gap that is widening and only ~25% attributable to peer-review/publication delay), with most abstracts generalising to “AI” rather than the model and configuration actually tested. The reason it matters is that governance imports its capability evidence from this literature; clinicians/lawyers/journalists/teachers/policymakers (who lack the knowledge and/or time to stay up-to-date with arXiv/METR/Epoch) read the academic record and the headlines downstream of it, so the literature’s staleness propagates into consequential decision-making in high-stakes applied domains. The audit is paired with policy recommendations, VERSIO-AI (a proposed academic reporting standard to mitigate this problem), and a per-DOI audit tool (frontierlag.org).
Pre-printarXiv:2605.041352026

Panels

Evals-Consensus.AI
Served on the expert Delphi panel, rating proposed practices across successive rounds alongside participants from leading labs (e.g. DeepMind, OpenAI) and orgs (e.g. AISI, METR, RAND); endorsed the resulting AI Evaluation Consensus Statement. My Safety Under Scaffolding paper was among the evidence base reviewed.
Expert panelist & endorser2026

Publications

Peer-reviewed

  • Kam, M … Gringras, D, et al. (2025). Delay to invasive coronary angiography for patients with NSTEMI admitted to hospitals without cardiac catheterisation facilities in South-East Scotland. British Journal of Cardiology, 32:63–7.
  • Gringras, D. (2024). From clinical to human connection: A reflective journey to patient-centred care. Lifestyle Medicine, 5(2).
  • McCallum, AK & Gringras, D, et al. (2023). Supporting drug and alcohol service users to cut down or stop smoking: review and evidence gaps. European Journal of Public Health, 33(2).

Pre-prints · several submitted for peer review

  • Gringras, D. (2026). Safety under scaffolding: How evaluation conditions shape measured safety. arXiv:2603.10044
  • Gringras, D & Salahshoor, M. (2026). Frontier Lag: A Bibliometric Audit of Capability Misrepresentation in Academic AI Evaluation. arXiv:2605.04135
  • Gringras, D. (2026). IatroBench: Pre-registered evidence of iatrogenic harm from AI safety measures. arXiv:2604.07709
  • Thwing, T, Gringras, D, Richard, U, Tang, M. (2026). Strengthening the International Network for Advanced AI Measurement, Evaluation and Science to Support Enforceable Global AI Red Lines. SSRN 6854278.
  • Gringras, D. (2026). Defensive AI: When safety alignment creates tort liability for medical information omission. SSRN — Artificial Intelligence: Law, Policy, & Ethics eJournal.

Manuscripts in peer review

  • Gringras, D, et al. (2025). Assessing UK Medical Ethics Education: Perspectives and Proficiencies of Graduating Medical Students. In peer review, BMC & BMJ journals.
  • Gringras, D, et al. (2025). Lifestyle Medicine: Perspectives and Proficiencies of Near-Graduate UK Medical Students. In peer review, BMC & BMJ journals.
  • Gringras, D, et al. (2025). A National Evaluation of Public Health and Health Policy Literacy Among UK Medical Students. Accepted at BMC Public Health — pending publication.