MSc Artificial Intelligence research · University of West London
Ask a UK annual report a question.
Get the tagged figure, not a guess.
Upload a Companies House iXBRL filing and ask about it in plain English. Each answer comes from the filing's own tagged XBRL facts. It shows its source concept, period and dimension, and whether the server could confirm it against the filing.
- Correct answers
- 44% → 92%
- Unanswered
- 52% → 3%
- Wrong when answering
- 5.1% vs 18.2%
Naive reading vs tag-aware retrieval, 102-question benchmark
Retrieval gives the model the evidence that naive reading loses to truncation
Claude vs GPT-5.6 Terra, with identical retrieved evidence
2 · Ask
Answers appear here as evidence cards.
- Matched to tagged fact value read directly from the filing
- Computed from tagged facts arithmetic done by the server
- Quote found in filing narrative answer backed by the tagged text
- Not verified the server could not confirm it
How it works
Parse
The iXBRL document is parsed directly with lxml. Scale, sign, nil values, continuation chains and
ix:excludeare resolved, and no taxonomy download is needed.Store
Every fact is stored with its concept, value, unit, period and dimensional qualifiers. The multiple values a filing tags for one concept are all kept.
Retrieve
Tag-aware retrieval selects the facts whose concepts match the question. The model reads structured evidence instead of a truncated slice of a 300,000-token report.
Verify
The model must cite evidence IDs. The server reads the value from the cited fact, does any arithmetic itself, and checks narrative quotes against the tagged text.
Research findings
| Condition | Correct | Incorrect | Abstained |
|---|---|---|---|
| Claude · C1 naive (8k-token text) | 44.1% | 4 | 53 |
| Claude · C3 tag-aware retrieval | 92.2% | 5 | 3 |
| GPT-5.6 Terra · C1 naive | 41.2% | 5 | 55 |
| GPT-5.6 Terra · C3 retrieval | 79.4% | 18 | 3 |
Retrieval fixes coverage. Both models stopped abstaining once they were given the relevant tagged facts. Claude's C1→C3 gain was +48.0 pp (95% CI 38.2–57.8).
Reliability depends on the model. With byte-identical evidence, GPT-5.6 Terra was wrong on 18.2% of the questions it attempted, against 5.1% for Claude.
Self-consistency did not detect errors. Sampling-based detection caught 0 of the incorrect answers for either model. The errors were systematic, so they did not vary between samples.
Limitations: a small corpus, three consistency samples, and results that do not transfer automatically to other filings. This demo uses a structured-output variant of the C3 prompt with server-side verification, so it is not the exact benchmarked condition.
Want more questions, or Claude access?
The public model is free for 10 questions per session. For more questions or a Claude access code, send me a message with a line about what you'd like to test.