Assessed 19 August 2026 against TRACE v0.1. Updated when the score changes.
Level 1 — Traceable
Every number shown carries a resolvable source, a retrieval date and an evidence level.
| Check | Result | What it means |
|---|---|---|
| Every number names its source | partial | Values retrieved through OPTIMADE carry their entry ID. Values that arrive from the model’s own knowledge rather than from a retrieval do not, and are not currently distinguished from those that do. This is the most important open defect. |
| The source resolves | pass | Entry IDs link out to the source database. |
| Source identifies the value, not just the material | partial | The link lands on the entry. Whether the specific property is visible there varies by database. |
| Origin distinguishable | partial | Computed values from DFT databases are labelled. Measured literature values are not consistently separated from computed ones, and predicted values are not yet a distinct category in the interface. |
| Retrieval date shown | fail | Not implemented. Until it ships, no value we produce can exceed E3 under the date rule — including values that would otherwise be clean E2. |
| Method and condition with mechanical properties | partial | Temperature is usually present. Product form and heat treatment condition often are not. |
| Units always present | pass | |
| Uncertainty or range shown | fail | Single figures are presented for properties that legitimately vary with treatment. |
Not achieved.
Level 2 — Auditable
| Check | Result | What it means |
|---|---|---|
| Exports a record, not a transcript | in progress | Report export is a pre-launch requirement. Until it emits a valid record it is a formatted transcript, which is not the same thing. |
| Rejected candidates recorded | fail | We present what was found, not what was discarded. |
| Weakest-link rule enforced | fail | Not implemented. |
| Constraints testable as recorded | partial | Free-text requirements are accepted and are not always converted into testable constraints before screening. |
| Limitations stated | fail | Not implemented. |
| Accountability named | fail | Reports do not name an accountable engineer. |
| Qualitative requirements kept qualitative | pass | By omission rather than by design — we do not score them because we do not handle them. |
Level 3 — Reproducible
| Check | Result |
|---|---|
| Same question, same answer | fail — generation is not constrained to determinism; repeated runs vary |
| Sources date-pinned | fail — live queries upstream, no snapshotting |
| Tool and model versions recorded | fail |
| Non-determinism declared | fail — not stated in the interface |
Coverage, which is a separate problem
Conformance is about how we handle a value. Coverage is about whether we have one at all, and ours is uneven in ways worth knowing before you sign up rather than after. The full picture is on the data page; the short version:
- Commercial grade designations bridge incompletely to composition-indexed databases. Ask for a specification minimum for a named grade and you may get a composition match instead of an answer.
- Formulated materials — polymers, elastomers, composites, concretes, coatings, adhesives — are outside the data layer entirely.
- Organic chemistry and life sciences — effectively zero coverage.
- Strongest on inorganic crystalline materials: battery and energy storage, catalysts, semiconductors, ceramics, metallic systems at the composition level.
None of this stops the research agent being useful inside its coverage. It does mean you should check the coverage before you check the price.
The published grade pages have been counted against the same standard and the result is on one page: 191 sourced values, 82 named gaps, and not a single value above E3, because both of the standards our figures cite are paywalled and unopened. That count is the most useful thing we can say about our own data, and it is not a good number.
Why publish this
Two reasons, and it is worth being direct about the second.
It is a commitment device. A public score we have to keep updating is harder to quietly abandon than an internal backlog, and the items above are now checkable by anyone deciding whether to trust the product.
It is also marketing. A vendor publishing its own failures reads as credible, and we are aware of that. The defence against it being merely a posture is structural: the checklist is vendor-neutral, the licence lets competitors run it against us and ship something that scores better, and the failures above are specific enough to verify.
If you run the audit on us and get a better result than we published, that is a bug in this page and we would like to hear about it.
What this page does not cover
- A timeline. Some of these are weeks of work and some are architectural. Publishing dates we then miss would defeat the purpose of the page.
- Accuracy. Conformance is about traceability, not correctness. We have not published an accuracy benchmark; when we do it will include the cases we lose.
- Competitors’ scores. Run the checklist yourself. We are not going to grade our competition, because a vendor scoring rivals with its own framework is exactly the thing this page exists to avoid.
Change log
| Date | Change |
|---|---|
| 19 Aug 2026 | First assessment, TRACE v0.1 |