iXBRL Grounded Q&A

MSc Artificial Intelligence research · University of West London

Ask a UK annual report a question.
Get the tagged figure, not a guess.

Upload a Companies House iXBRL filing and ask about it in plain English. Each answer comes from the filing's own tagged XBRL facts. It shows its source concept, period and dimension, and whether the server could confirm it against the filing.

Correct answers
44% → 92%

Naive reading vs tag-aware retrieval, 102-question benchmark

Unanswered
52% → 3%

Retrieval gives the model the evidence that naive reading loses to truncation

Wrong when answering
5.1% vs 18.2%

Claude vs GPT-5.6 Terra, with identical retrieved evidence

2 · Ask

Answers appear here as evidence cards.

  • Matched to tagged fact value read directly from the filing
  • Computed from tagged facts arithmetic done by the server
  • Quote found in filing narrative answer backed by the tagged text
  • Not verified the server could not confirm it

How it works

  1. Parse

    The iXBRL document is parsed directly with lxml. Scale, sign, nil values, continuation chains and ix:exclude are resolved, and no taxonomy download is needed.

  2. Store

    Every fact is stored with its concept, value, unit, period and dimensional qualifiers. The multiple values a filing tags for one concept are all kept.

  3. Retrieve

    Tag-aware retrieval selects the facts whose concepts match the question. The model reads structured evidence instead of a truncated slice of a 300,000-token report.

  4. Verify

    The model must cite evidence IDs. The server reads the value from the cited fact, does any arithmetic itself, and checks narrative quotes against the tagged text.

Research findings

102-question benchmark · 12 UK filings (8 FTSE 350 IFRS, 4 FRS 102)
ConditionCorrectIncorrectAbstained
Claude · C1 naive (8k-token text)44.1%453
Claude · C3 tag-aware retrieval92.2%53
GPT-5.6 Terra · C1 naive41.2%555
GPT-5.6 Terra · C3 retrieval79.4%183

Retrieval fixes coverage. Both models stopped abstaining once they were given the relevant tagged facts. Claude's C1→C3 gain was +48.0 pp (95% CI 38.2–57.8).

Reliability depends on the model. With byte-identical evidence, GPT-5.6 Terra was wrong on 18.2% of the questions it attempted, against 5.1% for Claude.

Self-consistency did not detect errors. Sampling-based detection caught 0 of the incorrect answers for either model. The errors were systematic, so they did not vary between samples.

Limitations: a small corpus, three consistency samples, and results that do not transfer automatically to other filings. This demo uses a structured-output variant of the C3 prompt with server-side verification, so it is not the exact benchmarked condition.

Want more questions, or Claude access?

The public model is free for 10 questions per session. For more questions or a Claude access code, send me a message with a line about what you'd like to test.

Unlock Claude Sonnet 4.6

Claude had the lowest error rate in the study, so access is by code. Enter the code you were sent.

No code? Request access

You've used your free questions

Thanks for trying the demo. For more questions or a Claude access code, message me and I'll send one.