Evidence, not promises
Published benchmark runs — numbers unedited.
Every figure on this page is unedited solver output from the published RQP sample reports — the same certified PDFs you can download in the Resource Room. The flattering numbers and the boring ones, including the run where our own product recommended staying classical.
Portfolio Optimization RQP · published sample run
What an honest quantum benchmark looks like
78 assets, aggregated into 23 binary decisions (17 single names + 6 thematic bundles), 3 sector-budget bands, QAOA depth p=4 — every one of the 8.4 million possible portfolios evaluated exactly, against a certified classical referee.
8,388,608
Portfolios evaluated — every single one
Exact 2²³ state readout, no sampling noise
0.0%
QAOA gap to certified optimum
Best QAOA portfolio = SCIP-proven optimum, bit for bit
+3.2%
Classical heuristic gap
465 heuristic candidates — none reached the optimum
137×
Odds boost for the optimal state
vs uniform sampling — yet still 0.0016%, and the report says so
Winning portfolio: 8.0% expected return · 15.3% volatility · Sharpe 0.33 · budget gap 0.08% Run: 2 h 28 min certified end-to-end · warm-started p=4 · statevector simulator · seed 42, reproducible
- ✓
Verified, not asserted. QAOA's best portfolio equals the referee's proven optimum — gap 0.0%, certified by exact branch & bound (MIQP/SCIP) — while the classical heuristic fell 3.2% short.
- ✓
No magic claimed. The optimum's sampling probability was 0.0016% — a 137× boost over uniform, and still a small number. Both figures are printed in the report.
- ✓
Scale is stated, not implied. Simulator memory doubles with every qubit — around 30 is where real hardware has to take over. Prototyping below that line, with a referee, is how you build a quantum case you can defend.
Index Replication RQP · published sample run
Your mandate rules have a price. The quantum benchmark certifies it.
Enhanced-indexing the S&P 500 with 20 instruments: 500 assets in the objective, 17 selection qubits — and a referee that re-solves the mandate to price every rule in tracking error.
4.473%
Certified minimum tracking error
SCIP dual bound proves no basket beats it · MIP gap 0
0.00%
QAOA gap to certified optimum
Its basket is a certified-optimal one
+3.39%
Greedy heuristic gap
The industry-standard shortcut missed; random baseline +1.20%
The constraints landscape — every rule priced by the referee
Certified cost = optimum re-solved with only that rule removed. Unedited from the report.
Technology band
4–6 names required · landed at 6
+0.332%
Climate-transition band
2–4 names required · landed at 4
+0.339%
Both rules together
the package overlaps — it costs no more than its strictest rule
+0.339%
The irreducible floor: the instrument pool and exclusions alone impose 3.098% tracking error before any mandate rule applies.
Run: 110 s end-to-end · QAOA 53.8 s (p=2, 3 restarts, exact readout) · referee proof 11.2 s · 17 qubits, 544 two-qubit gates
- ✓
Constraints stop being opinions. Each mandate rule is priced by re-solving the certified optimum with only that rule removed. The mandate discussion becomes numbers, not beliefs.
- ✓
The irreducible floor is stated. The instrument pool and exclusions alone impose 3.098% tracking error before any rule applies — whatever the pool cannot replicate shows up openly, not hidden in a heuristic.
- ✓
No magic claimed. Only 3.6% of raw quantum readout hit the 20-name cardinality; candidates were repaired, and the best one reached the optimum after a single disclosed polish swap. On today's hardware this 544-two-qubit-gate circuit would retain ~7% fidelity. It's all printed in the report.
Robust Stochastic Collateral Allocation RQP · published sample run
The cheapest allocation and the stable one are different. The referee prices both.
Tuesday-morning margin calls at a Swiss bank treasury: three venues calling at once, 23 collateral lines — ten locked, three pledge pools that move whole — 50 stress scenarios, and a 24-variable register. A scenario-covariance term trades posting cost against coverage stability, and an exact referee enumerates every valid allocation to prove the optimum.
16,777,216
Allocations evaluated — every single one
Exact 2²⁴ state readout, no sampling noise
0.00%
QAOA gap to certified optimum
Best sampled allocation = the enumeration-proven optimum
2.6%
The optimum is the modal state
The single most likely readout of the whole register — ~428,000× uniform
CHF 19k/yr
The certified price of coverage stability
Frontier flip: posting cost that buys coverage volatility from 8.3M down to 7.6M (flagship run)
Winning allocation: govt-bond ladder to LCH · Pfandbrief pool to Eurex · equity basket to the CSA — equities where they are cheapest to give away, priced at a 14% scenario undercoverage probability the stress panel prints openly Run: 29 min end-to-end · warm-started CVaR QAOA p=3 · exact readout of all 16.8M states · fixed seeds, reproducible
- ✓
Seeded, and it says so. The register starts warm — rotated toward the best classically-enumerated allocation — and the report states it. The quantum sampler's job is to concentrate on or improve that incumbent; the exhaustive referee certifies whichever answer any method produces. On this run it concentrated so hard the optimum became the single most likely state.
- ✓
The first runs failed, and it's printed. Expected-energy training sampled an infeasible best allocation, and a too-coarse slack encoding once made the QUBO's own minimum provably infeasible — no optimizer could have sampled its way out. Both failures are published; the auto-calibration they forced now prints its choices in every report.
- ✓
Stress is classical, and labeled. The coverage-under-stress panel and the cost-vs-stability frontier are classical scenario analysis over the workbook's own scenarios, labeled as such in the report. The quantum content is the allocation decision; the tail machinery's quantum counterpart — amplitude estimation — lives in the QMC Lab, measured transparently.
- ✓
Above the cap, quantum steps aside. The full-scale flagship workbook (36 QUBO variables) is beyond the statevector simulator, so its quantum stage is skipped by design — the run completes classically with the certificate, frontier, and stress panels at full scale, and the report says exactly that.
QML Classification RQP · published sample runs
Don’t bet on quantum. Measure it.
Classifying German power-price spikes with an 8-qubit quantum kernel against six tuned classical baselines — 17,157 rows, an exact 10,294×10,294 quantum kernel, paired bootstrap statistics. For this dataset the measured answer was “not yet” — and that verdict is the product working as designed. The counterpart below — a quantum-native demo dataset — passed the same gate and finished with the whole confidence interval above 1.0.
34/100
QML suitability score — this dataset
“Too simple for quantum” — every dataset is scored before a single circuit runs (±12)
0.92×
Advantage factor, raw features
95% CI 0.896–0.937, entirely below 1.0 — classical clearly ahead
0.99×
Advantage factor, Fourier-engineered features
95% CI 0.975–1.011, includes 1.0 — “inconclusive within noise”
1 day
The cost of knowing
Verdict for this dataset, today: keep the classical model — re-tested as encodings and hardware evolve
How these verdicts are produced
Two assessment levels gate every training run: Fast (seconds — classical probes and a suitability score) and Deep (minutes — kernel diagnostics and the five-condition checklist). Methodology after the Fourier Wall paper (arXiv:2607.15815); every check prints in the sample reports.
The same gate, passed — quantum-native demo dataset (26 Jul 2026)
1.45×
Advantage factor, quantum-native demo
95% CI 1.31–1.62, entirely above 1.0 — a certified win against the full classical bar, order-matched twin included
0.89
VQC test F1 — vs 0.71 best classical
6-qubit variational classifier on the quantum-teacher demo dataset — full report and dataset in the Resource Room
The deltas, audited — complementary-signal panel, BAF fraud benchmark (7 Aug 2026)
0.00
Quantum blend weight — fitted on validation
The best classical pairing declines the quantum score entirely (BAF fraud benchmark, fraud-ops configuration, 15-qubit kernel)
+0.0002
Ensemble delta vs classical-stack control
95% CI −0.0004 to +0.0011 — statistical zero. The control single-model comparisons omit: a stack of every classical baseline
0.92×
Advantage factor, fraud-tuned run
95% CI 0.83–1.00 — classical ahead. Printed verdict: no complementary signal
- ✓
The bar is real. Quantum is scored against tuned classical baselines — including an order-matched periodic twin that reads, up to interaction order 3, exactly the structure an angle-encoded quantum model could pass off as “advantage” (methodology after the Fourier Wall paper, arXiv:2607.15815).
- ✓
Statistics, not vibes. An advantage is claimed only when the whole 95% confidence interval sits above 1.0. Better feature encoding moved the factor from 0.92 to 0.99 — a measured trajectory, not a promise. Today’s printed verdict: inconclusive within noise — and the five-condition quantum-territory checklist closes the case: still classical territory.
- ✓
Even the deltas get audited. The popular ensemble claim — “the quantum model sees what the classical models miss” — is now tested on every binary run: quantum-vs-classical disagreements are scored against a classical-ensemble control on identical rows. On the public BAF fraud benchmark, under a fraud-ops configuration, the fitted quantum blend weight came out 0.00 and the ensemble delta sat within noise — printed verdict: no complementary signal. Lift over a single classical model is ordinary ensemble diversity; the panel makes that distinction measurable.
- ✓
A verdict about one dataset — not about quantum. Other data profiles score differently; that is what the suitability filter is for. And “let's revisit later” only works if someone is measuring: the harness re-runs as encodings and hardware evolve, so the week the interval clears 1.0, you know.
Sources & scope
Portfolio run: published sample report of 16 Jul 2026 (78-asset demo universe). Index run: published sample report of 12 Jul 2026 (S&P 500 demo mandate, 20 ISINs). QML runs: published sample reports of 23 Jul 2026, re-verified 26 Jul 2026 under the upgraded classical bar (German power-price spike demo dataset, baseline & Fourier-engineered). QML assessment methodology: independent implementation inspired by Javier Mancilla and Tomas Tagliani, “The Fourier Wall: Why Public Tabular Datasets Refuse Quantum Advantage, and a Certified Recipe for Where It Lives” (arXiv:2607.15815); qubit-lab.ch is not affiliated with the authors. All reports are gated downloads in the RQP Resource Room. All figures are unedited solver output on demo data — illustrative only, not investment advice, not indicative of results on other problems.
Let’s explore — no strings attached
A relaxed 30-minute call or a live demo.
Management-friendly or physicist to physicist — your data, your questions, real runs. No slides, no obligations.