What a New AI Benchmark Gets Right About NEPA Drafting — And What It Misses
A research team at Pacific Northwest National Laboratory just published the most rigorous evaluation to date of AI systems for NEPA document drafting. The paper, DraftNEPABench, was published as a PNNL/OpenAI preprint and represents a serious, methodologically careful effort to measure how well current AI coding agents can draft Environmental Impact Statement sections.
The findings are worth understanding carefully — both for what they confirm and for the gaps they reveal.
What the benchmark actually tested
DraftNEPABench comprises 102 EIS drafting tasks drawn from 22 real environmental impact statements, spanning 18 lead agencies and 18 distinct EIS section types — from biological resources and water quality to socioeconomics and visual resources. Each task was created by subject matter experts with real EIS drafting experience, and drafts were evaluated on four dimensions: accuracy, clarity, reference use, and structure.
Three major AI coding agents were tested: Claude Code (Anthropic), Codex CLI (OpenAI), and Gemini CLI (Google). These were compared against standard retrieval-augmented generation (RAG) baselines using the same underlying models.
The headline finding: agents win on structure, lose on facts
Coding agents substantially outperformed RAG baselines across every judge and every evaluation criterion. Claude Code ranked first overall, receiving an average SME score of 4.10 out of 5 — with reviewers noting strong document structure, logical organization, clear transitions, and effective reference use. In several cases, SMEs judged Claude Code outputs as comparable to or better than the human-written ground truth drafts.
But the paper's most important finding is also its most sobering: factual accuracy remained the persistent bottleneck across every system tested, including the best-performing agents. SME reviewers consistently flagged issues with factual specificity, data placement, and citation correctness. The paper states directly that "improved drafting workflows alone do not guarantee factual grounding."
This distinction matters enormously in permitting practice. A well-structured document with incorrect regulatory citations or imprecise project-specific data doesn't just get marked down in a benchmark — it triggers agency requests for information, delays review timelines, and in some cases undermines project credibility with the lead agency entirely.
The counterintuitive difficulty finding
One underappreciated result deserves attention. The paper found no performance penalty on harder EIS sections — those rated as most demanding by human experts held up just as well as the easier ones, contrary to what most would expect. The researchers hypothesize that more complex sections tend to have richer explicit structure and clearer constraints, which benefits agentic systems.
For energy project developers and environmental consultants working on highly structured permit applications — like the nine-attachment Class VI UIC permit package required for carbon storage projects — this is relevant. Structured, constraint-rich documents may actually be better candidates for AI-assisted drafting than the more open-ended narrative sections where agents currently underperform.
What the benchmark doesn't cover
DraftNEPABench is a meaningful contribution, but its coverage reflects the established EIS universe rather than the frontier of energy transition permitting. Several critical gaps stand out:
Class VI UIC permitting under the Safe Drinking Water Act is entirely absent. CCS-specific review pathways — including EPA Underground Injection Control requirements, area of review analysis, and post-injection site care — are not represented. State-level environmental review processes, which increasingly run parallel to federal NEPA review for large infrastructure projects, are also outside the benchmark's scope. And two agencies that are structurally central to offshore wind, interstate pipelines, and energy export infrastructure — FERC and BOEM — are not meaningfully represented in the dataset.
These aren't minor omissions. For developers navigating carbon capture, offshore wind, or large transmission projects, the permitting complexity that determines project timelines and capital risk lives precisely in these gaps.
What this means for permitting practice
The PNNL benchmark confirms something practitioners already know: AI can meaningfully accelerate the structural and organizational work of environmental document preparation. What it cannot yet do reliably is substitute for deep regulatory grounding — the kind that comes from understanding how a specific agency has interpreted specific requirements for similar projects in similar geographies.
That gap between clean drafts and defensible drafts is where permitting risk actually lives. It's also where Permica is focused.
We're building regulatory intelligence specifically for the permitting challenges that general-purpose AI benchmarks don't yet reach — starting with Class VI carbon storage permitting, where the regulatory corpus is dense, the agency precedent is still forming, and the cost of a poorly grounded application is measured in years, not weeks.