Introducing Reg Affairs Bench

AI assistants answer regulatory questions well, and the deliverables behind those questions are harder. A change control must be classified against the guideline in force at submission. A Module 3 must agree with itself across a dozen documents. A memo must say plainly when a validation report is missing. One wrong reporting category means a deviation, a delayed launch or an agency query.
Existing benchmarks cover neighbouring ground. FDARxBench1 tests document-grounded questions on FDA drug labels. LifeSciBench2 tests life science research workflows. Harvey's Legal Agent Benchmark3 tests long-horizon agent work in law. We created Reg Affairs Bench to test the deliverables regulatory teams actually produce: change classifications, dossier checks and submission drafts. We used the benchmark to compare Arca with frontier model APIs on the same models.
- Tasks in the latest headline set
- 33
- Model and harness combinations
- 4
- New tasks awaiting expert review
- 7
- Regulatory affairs experts writing and reviewing earlier rubrics
- 4
The tasks
Each task reads like a request to a senior colleague: a scenario, its documents and the deliverable expected back. Answers are free text. A task-specific rubric awards points for the decisions, citations and open questions an expert expects. A fail line zeroes any answer that makes a disqualifying call, such as agreeing to a CBE-30 when the change needs prior approval.
Classify a post-approval change: EU variation type and code under the guideline in force at submission, or the US category (prior approval supplement, CBE-30, CBE-0, annual report). Credit for checking each condition against the facts, grouping related changes, naming impacted dossier sections, and asking the questions that could change the outcome.
Audit a set of Module 2 and 3 documents for conflicts. Real conflicts are planted next to differences that are supposed to exist; answers lose credit for false alarms as well as misses.
Draft a dossier section when a required report is not in the package. Credit for naming the missing document and what it blocks; fail lines for inventing its results or implying agency agreement that never happened.
Everyday regulatory deliverables: agency correspondence, label reviews and clinical hold responses.
Find every document in an 838-file product library that states a value, or the one document that decides an answer, with ground truth known by construction.
Construction
Ground truth. EU variation tasks start from published EMA procedural records. We rewrote each as a change control for a fictional company and wrote the dossier documents it needs. The published classification stays the answer. US reporting-category tasks are checked line by line against 21 CFR 314.70, 601.12 and FDA's post-approval guidance.
Planted conflicts. Consistency tasks hide real conflicts, such as a 36-month shelf life in the Quality Overall Summary against 24 months in the stability conclusion. They also hold differences an expert should not flag, such as release and shelf-life limits. A leak check confirms that no document contains the answer or the rubric's wording.
Rubric authors. Three external regulatory affairs professionals and one contractor engaged by Arca wrote and reviewed the rubrics. Two of them wrote rubrics for the same ten tasks without seeing each other's work, and we merged those into one rubric. A third reviewed the earlier EU variation rubrics, and the contractor reviewed other tasks. The seven newly added EU tasks have model-drafted rubrics and await expert review; their results are provisional.
Model jury. Models from three vendors grade every criterion. The full jury steps in whenever a verdict is partial or a fail line might fire. Judges see the source documents, so a fact from the library is never mistaken for an invented one.
Example task
This example comes from an earlier evaluation run; it illustrates the rubric rather than the latest aggregate scores. A cell therapy site adds eight more instruments to bays it is already licensed for. Both systems recommend a prior approval supplement, so the gap between them is in the supporting work: which guidance row governs the change.
Sample task
Corvane Cell Therapeutics holds BLA 197388 for TESSACEL (tessagenlecleucel), an autologous anti-CD19 CAR-T product made one patient lot per AX-200 automated closed processing instrument in Suite 2B, Frederick, MD. Change control CC-2026-0512 adds eight more AX-200 instruments of the same model, software version, tubing set and locked recipe into the eight empty bays of the same licensed suite, doubling the lots that can be in process at once from 8 to 16. Process parameters, materials, specifications, room classification and HVAC are unchanged, and the suite was qualified for 16 bays. Write a US change assessment memo with the reporting category for each element and for the submission as a whole.
- 3.2.A.1 facilities and equipment
- 3.2.P.3.3 description of manufacturing process
- APS-2026-03 aseptic process simulation summary
- ER-2026-07 engineering run report
Reference answer
Prior approval supplement for the change as a whole. The change owner's CBE-30 rests on the guidance row for a duplicated process train, which does not cover doubling concurrent autologous lots in a room that has never run at full occupancy.
ArcaClaude Opus 5
- Score
- 100
- Points
- 100/100
- Steps
- 9
- Cost
- $1.26
- Time
- 8.3 min
Arca pulls the FDA guidance rows for capacity and for duplicated process trains, sees that the row the change owner cites does not fit a cell therapy product at higher occupancy, and files the whole change as a PAS.
Claude Opus 5 API
- Score
- 62
- Points
- 62/100
- Steps
- 6web searches
- Cost
- $1.36
- Time
- 5.9 min
The plain API lands on the same recommendation and argues it well from the regulation, but it never finds the two guidance rows the answer key asks for, so it loses the points for showing which row governs.
Final answer
Prior approval supplement (PAS) for the change as a whole. The proposed CBE-30 is not supportable, and the guidance line the change owner is relying on is the wrong row for a cellular therapy product.
Final answer
...the Appendix line the change owner cites supports the instruments; it does not support the doubling of concurrent autologous lots in a single room that has never been run at that occupancy. Either file the PAS, or register only what has been demonstrated and earn the rest with a full-occupancy simulation campaign.
Trajectory
- FDA change lookupaddition of duplicated process train or process step with no change in process parameters (BLA)
- FDA change lookupincrease in manufacturing capacity, scale up, number of production lines, additional equipment (BLA)
- FDA change lookupfacility change, addition of manufacturing area within licensed building, environmental monitoring (BLA)
- FDA change lookup3.2.P.3.3 addition of identical unit processes, concurrent manufacturing throughput, cell therapy (BLA)
- Legal search21 CFR 601.12 changes to an approved application, biologics reporting categories
- Fetchfda.gov, 2021 CMC changes guidance for certain biological products
- FDA change lookupenvironmental monitoring program, addition of sampling locations, aseptic processing area (BLA)
- FDA change lookupmedia fill, aseptic process simulation requalification, closed system processing (BLA)
- AnswerMemo returned
Trajectory
- 6 web searchesUsed the model's built-in web search
- AnswerMemo returned
Rubric
Rubric
Validating the benchmark
Two regulatory affairs experts wrote rubrics for the same ten tasks independently. Scoring the same 69 answers with each rubric gives a Pearson correlation of 0.86. The seven systems rank almost identically under either rubric (Spearman 0.96). Where the experts disagreed on substance, the merged rubric follows the written guidance and records why.
A single run can move one task by twenty points. The latest headline uses one saved answer per task and system, rather than the earlier averages over repeated runs. The expert-agreement study above is separate from this updated run.
Results
Updated 28 September 2026: these preliminary results cover 33 tasks. We reused the saved answers for 26 unchanged tasks and ran seven new tasks on each system. The seven new tasks are based on published EMA filings where a model recommended a heavier reporting category than EMA accepted. Their answer keys await expert review.
Lift over the plain APIs
- Claude Opus 5
- +11.5
- Arca 88.2 vs plain API 76.7
- 33 tasks, one answer per task
- GPT-6 Astra
- +4.6
- Arca 85.6 vs plain API 81.0
- 33 tasks, one answer per task
| Model | Arca | Plain API | Arca vs API |
|---|---|---|---|
| Claude Opus 5 | 88.2 | 76.7 | +11.5 |
| GPT-6 Astra | 85.6 | 81.0 | +4.6 |
The plain APIs include their native web-search tools, with the task documents supplied in context. These are comparisons against search-enabled APIs, not models answering from memory alone. The open-source agent and DeepSeek have not been evaluated on this updated set, so their earlier scores are not included here.
Making the set harder
We retired two consistency tasks and both retrieval tasks from the headline set, and added seven EMA-based over-filing tasks. The retrieval tasks remain available separately: systems with a document index scored 94–100 on them.
| System | Previous set, single-run baseline | Updated 33-task set |
|---|---|---|
| Arca + Claude Opus 5 | 90.1 | 88.2 |
| Arca + GPT-6 Astra | 88.6 | 85.6 |
| Plain GPT-6 Astra API | 85.0 | 81.0 |
| Plain Claude Opus 5 API | 78.6 | 76.7 |
These previous-set baselines use the canonical single runs. They differ from the repeated-run averages in the original article.
The added tasks average 74–86 across systems. Each model makes mistakes on different tasks, so adding them lowered the headline scores by only 2–4 points. We have not yet reached the difficulty we want.
Open problems
- Expert review. The seven new answer keys need review. An EMA outcome is evidence for a real filing, but a reconstructed scenario must contain enough facts to justify that outcome.
- Judgment calls. On one cell therapy change, EMA accepted Type IB. The rubric allows partial credit for Type II when the recommendation rests on a comparability gap. We need to distinguish model mistakes from reasonable disagreements with the key.
- Broader comparisons. Consumer apps and the open-source agent need runs on the updated set before we can compare them with these headline results.
Scope and next steps
The current headline covers 33 regulatory affairs tasks. Its rubrics combine expert-written or reviewed keys with provisional, model-drafted keys. Scores measure the written deliverable; they do not measure time saved or the outcome of a real submission. The earlier cost chart and fail-line rates have been removed because they were not recalculated for this set.
Next we are asking 20 experts for difficult cases from their past work, including situations another specialist could get wrong. We will test these with Arca and ChatGPT, check whether low scores reflect substantive errors, and commission more cases and detailed rubrics from the most promising contributors. We are also exploring tasks that require specialist data sources and evidence spread across a full Veeva document library.
References
- Xiong et al. (2026), "FDARxBench: Benchmarking Regulatory and Clinical Reasoning on FDA Generic Drug Assessment." arXiv:2603.19539↩
- OpenAI, "Introducing LifeSciBench." openai.com↩
- Harvey, "Introducing Harvey's Legal Agent Benchmark," 6 May 2026. harvey.ai↩
