We are early access, and everything here runs on public or simulated data — each result traces to a dated report and names the data it came from. Where our own tests favour the simpler model, we say so.
On stale or incomplete data the system declines to predict rather than guessing.
In the study the gate rejected about 70% of rows outright and issued no forecast on them. On the rows it accepted, error fell between roughly 1.5× and 10× depending on how the data degraded — 10× is the worst corner tested, not the typical case. The rejection rate is the point rather than a caveat: the value sits in the rows it refuses.
Two limits worth stating. This is a single sample — gap positions were randomly drawn, so the direction is robust and the exact numbers move. And it is a claim about the gate, not about our predictor: a gradient-boosted baseline scored marginally better than ours on the accepted rows. We describe this as knowing when not to predict, because that is what the result actually shows.
It is algorithm-agnostic and hard to copy, and it is the behaviour an operator actually needs from a system they are going to trust.
Only the pair is informative. A page that shows one without the other is marketing.
A prospect who finds these after we have hidden them concludes we hide things. A prospect who finds them because we published them concludes our numbers can be trusted.
On real commercial cells from the MIT/Stanford dataset, the plain data-driven model placed first and our full physics model placed fourth. Reported exactly as found.
On the IGBT reproduction, the physics-informed model underperformed a standard LSTM across the folds. We marked it unstable and archived it rather than quietly dropping it.
On the NASA battery ablation, the improvement came from the temperature-aware formulation rather than the physics terms. We corrected our own attribution rather than let the wrong story stand.
Across four independent evaluations the evidence does not support a blanket claim that physics improves accuracy. What it supports is rarer and more valuable: a company that scores itself honestly, publishes its misses, and can trace every number to a dated source.
Accuracy is one question; whether it runs on a box at the site is another. Four hours of concurrent inference on an NVIDIA Orin NX, measured.
Measured on an NVIDIA Orin NX during an edge-deployment benchmark, May 2026 (SOW-EDGE-5). These are deployability measurements — latency, endurance and GPU headroom — not accuracy claims. CPU ran near its ceiling at peak, so how many assets one box can carry is a separate validation.
A read-only assessment proves the forecast on your own data before any hardware conversation. You see the record; you decide what it is worth.