Making an LLM verifiable on French accounting law: what we measured
We built a dated corpus, a tool-calling agent and a benchmark for French accounting law. Citations that resolve: 98–100% with the harness, 57–75% without.
Blog
Welcome to Nodal’s blog. We write about AI, cybersecurity and software engineering, analyzing the ongoing industrialization of intelligence in terms of its impact on how systems are built, maintained and protected.
We built a dated corpus, a tool-calling agent and a benchmark for French accounting law. Citations that resolve: 98–100% with the harness, 57–75% without.
Strip out the AI and the Hugging Face incident is an ordinary operational failure: an uninventoried shared resource, a wrong escalation bar, guardrails that blocked defenders.
OpenAI's postmortem names four patterns of misaligned behaviour. METR's independent investigation tells a stranger story — and the two accounts don't fully agree.
In July 2026, OpenAI agents broke out of a benchmark sandbox and reached Hugging Face's production clusters. The story, told once, in order.
A 27-year-old OpenBSD flaw, a 16-year-old FFmpeg bug that survived five million fuzz runs, certificate forgery in wolfSSL: the concrete bugs behind “AI finds vulnerabilities”.
Anthropic's Mythos found 26,153 vulnerabilities; 202 are fixed. What Project Glasswing revealed about a disclosure system built around how fast humans find bugs.