Home — Evidence
Evidence

What we can prove, proven honestly.

We are early access, and everything here runs on public or simulated data — each result traces to a dated report and names the data it came from. Where our own tests favour the simpler model, we say so.

The proof we lead with

The strongest honest thing the system does is know when to stay quiet.

MEAN ABSOLUTE ERROR (Ah) — LOWER IS BETTER 0.1367 All rows — no gate 0.0135 Accepted rows — ~30% of rows ≈10× worst case
The gate rejected about 70% of rows. Error fell on the rest.Mean absolute error, all rows versus accepted rows, at the worst corner of the study — 30% of readings missing, in long blocked gaps. With isolated dropouts the same gate gives about 1.5×. Public NASA reference cells, missing-data robustness study, June 2026.

On stale or incomplete data the system declines to predict rather than guessing.

In the study the gate rejected about 70% of rows outright and issued no forecast on them. On the rows it accepted, error fell between roughly 1.5× and 10× depending on how the data degraded — 10× is the worst corner tested, not the typical case. The rejection rate is the point rather than a caveat: the value sits in the rows it refuses.

Two limits worth stating. This is a single sample — gap positions were randomly drawn, so the direction is robust and the exact numbers move. And it is a claim about the gate, not about our predictor: a gradient-boosted baseline scored marginally better than ours on the accepted rows. We describe this as knowing when not to predict, because that is what the result actually shows.

It is algorithm-agnostic and hard to copy, and it is the behaviour an operator actually needs from a system they are going to trust.


Forecast versus actual

The hit and the miss, side by side.

Only the pair is informative. A page that shows one without the other is marketing.

STATE OF HEALTH vs CYCLE — ACTUAL vs FORECAST 80% state of health — end of life forecast anchored here actual forecast ensemble spread end-of-life line
A forecast tracking to the end-of-life line.Redrawn from the PredictPowr™ evaluation view on public NASA reference cells. Ranges are ensemble spread, not calibrated confidence intervals.
THE MISS — THE MODEL NEVER CALLED THE CROSSING 80% state of health — end of life actual crosses… …forecast never does actual forecast
A cell where the forecast never crossed.The actual state of health fell below the end-of-life line and the forecast did not follow. Published rather than dropped.
B0005 Pass B0006 Pass B0007 Partial / late B0018 Partial B0030 No crossing B0042 Partial / late risk
Every retained reference cell, scored.Per-cell checkpoint summary across the retained public reference cells — passes, partials, and the one with no crossing.

We publish our misses

Three cases where the simpler model won — reported as found.

A prospect who finds these after we have hidden them concludes we hide things. A prospect who finds them because we published them concludes our numbers can be trusted.

MIT / Stanford · real cells
1st vs 4th

The plain model placed first.

On real commercial cells from the MIT/Stanford dataset, the plain data-driven model placed first and our full physics model placed fourth. Reported exactly as found.

Real commercial cells. Named because hiding it would cost more than the result.
IGBT reproduction
LSTM won

A standard LSTM beat our physics model.

On the IGBT reproduction, the physics-informed model underperformed a standard LSTM across the folds. We marked it unstable and archived it rather than quietly dropping it.

Independent reproduction study. Marked unstable, archived, disclosed.
NASA battery ablation
We corrected it

We corrected our own attribution.

On the NASA battery ablation, the improvement came from the temperature-aware formulation rather than the physics terms. We corrected our own attribution rather than let the wrong story stand.

Public NASA data. Attribution corrected on the record.
TEST MSE ×10⁻³ — LOWER IS BETTER Device 2Device 3Device 4Device 5 Plain LSTM Physics-informed LSTM
The IGBT reproduction, fold by fold.Leave-one-device-out cross-validation, IGBT remaining-life reproduction on NASA PCoE Dataset #8. The physics-informed model was worse on three of four devices. Published because it is a miss.

The scorecard

Which technique actually won.

Real commercial cells (MIT/Stanford) Best performer: Plain data-driven model physics did not win IGBT reproduction (NASA PCoE) Best performer: Standard LSTM physics did not win NASA battery ablation Best performer: Temperature-aware formulation physics did not win Simulated SiC devices Best performer: Neural model — the architecture, not the physics physics did not win
Four evaluations, reported as found.On the fourth — simulated SiC devices — the physics term made the model train faster, not predict better. Re-run at a correct training budget, the physics advantage was 0.6% and not statistically significant. The neural model did beat the gradient-boosted baseline on all four measures, on both test conditions, in every one of ten runs — that is the architecture, not the physics. The data is physics-based synthetic, so it speaks to method behaviour, not to measured devices. Margins stay under diligence until real-device validation exists.

Across four independent evaluations the evidence does not support a blanket claim that physics improves accuracy. What it supports is rarer and more valuable: a company that scores itself honestly, publishes its misses, and can trace every number to a dated source.


Deployability

And it runs where the asset is.

Accuracy is one question; whether it runs on a box at the site is another. Four hours of concurrent inference on an NVIDIA Orin NX, measured.

1.70 ms
p99 inference latency — battery remaining-capacity model.
0.31 ms
p99 inference latency — SiC surrogate model.
4 hours
Continuous concurrent operation, both models, no drift or intervention.
2% GPU
Peak GPU use — inference ran on the CPU, leaving the GPU idle for heavier models and higher-rate signal processing. Not headroom for more assets.

Measured on an NVIDIA Orin NX during an edge-deployment benchmark, May 2026 (SOW-EDGE-5). These are deployability measurements — latency, endurance and GPU headroom — not accuracy claims. CPU ran near its ceiling at peak, so how many assets one box can carry is a separate validation.

See the proof on your own assets.

A read-only assessment proves the forecast on your own data before any hardware conversation. You see the record; you decide what it is worth.

Software-first · no raw data leaves your network · every number traces to a dated report