Part of TRACE, CC BY 4.0. Run it on us too — our score.
Start with the thirty-second test
If you only have time for one check, do this one.
Ask the tool for a specific property of a specific material. Take the number it gives you. Try to reach the exact source of that exact number in under two minutes.
If you cannot, the value is E0 — whatever the interface looked like — and it does not belong in your documentation.
This is not a test that language models fail by construction. A model that hands you a citation you can open passes it. It is also not a test only AI tools take: run it on a search engine, on a handbook PDF, on a colleague’s recollection. E3 is entirely serviceable for screening. The point is knowing which level you are standing on when you sign.
The three levels
Level 1 — Traceable. Every number presented carries a resolvable source, a retrieval date and an evidence level. No number is shown without them.
Level 2 — Auditable. Level 1, plus the tool exports a record rather than a transcript, records what it rejected and why, enforces the evidence rules rather than suggesting them, and states its own limitations.
Level 3 — Reproducible. Level 2, plus sources are date-pinned, an identical re-run produces an identical record, and tool and model versions are recorded.
How to score it
Do not average. A tool that passes seven Level 1 checks and fails one is not “88% Level 1”. It is not Level 1. Evidence does not average.
Score the default path. If provenance exists but sits behind a setting nobody turns on, it fails. The question is what an engineer in a hurry gets, not what the tool can do when studied.
Free tools are in scope, and it is worth scoring one at least once. The point is not to disqualify a search engine — it is to see, in writing, what level your habitual sources actually sit at.
Run it
[interactive checklist component — nineteen checks, live scoring, export]
What to do with the result
If your main tool scores Level 1 or better, you are in reasonable shape. Record which level in your selection notes; the next reviewer will want to know.
If it scores below Level 1, that is not a reason to stop using it. Most engineering runs on E3 sources and always has. It is a reason to stop treating its output as documentation, and to write down where each number came from yourself.
If a value governs a design allowable, the level matters more than anywhere else. Below E5, it needs your explicit, recorded acceptance rather than an invisible default.
What this checklist does not do
- It does not measure accuracy. A tool can score Level 3 and be wrong. Traceability and correctness are different properties, and this measures the first.
- It does not compare feature sets. Nothing here scores usability, coverage, price or support, all of which matter and none of which are what this instrument is for.
- It does not account for coverage. A tool with perfect provenance and no data for your material scores well and helps you not at all. Check coverage separately — ours is here.
- It is not neutral about what matters. It is built by people who think provenance is the thing worth measuring. That is a position, not a fact, and you should weigh it as one.