Quantum Benchmark Zoo
A community-maintained catalog of protocols for measuring quantum computer performance, from single-gate characterization through whole-processor scores and application suites to error-corrected logical qubits, the software stack, and analog platforms. Each entry covers what the benchmark measures, how it works, key papers, and reference implementations.
All benchmarks
-
Scalable characterization protocol that estimates the Pauli error rates of every gate and measurement on a processor simultaneously from a small set of shallow Clifford circuits
-
IonQ's single-number application metric, the largest width at which a six-algorithm circuit suite clears a 1/e fidelity bar, retired in 2025 in favor of industry-standard metrics.
-
Estimates the many-body fidelity of an analog quantum simulator's quench dynamics from about a thousand bitstring samples, doing for Hamiltonian evolution what XEB does for random digital circuits.
-
Automated head-to-head comparison of quantum compilers (Qiskit, Tket, Cirq, Quilc, PyZX) on gate counts, depth, circuit cost, and runtime, with auto-generated LaTeX reports
-
MetriQs-France benchmark suite scoring whole quantum stacks on optimization, linear systems, many-body simulation, and factoring, aggregated into one user-weighted figure of merit.
-
Two-copy benchmark that measures pairs of identical circuit outputs in the Bell basis, yielding fidelity estimates and circuit diagnostics that stay efficient beyond the classically simulable regime.
-
IBM's open-source suite of 1,000+ timed tests comparing how fast quantum SDKs construct, manipulate, and transpile circuits, plus the quality of the compiled output.
-
Streamlined, fully scalable RB that drops motion reversal entirely: circuits of i.i.d. random layers act on a random Pauli eigenstate, and each shot returns a simple pass/fail outcome.
-
Uses group character theory to isolate clean exponential decays, extending randomized benchmarking to gate sets that form groups beyond the multi-qubit Cliffords.
-
Open-source Python suite that generates Ising, QUBO, and higher-order optimization problems with planted, a-priori-known ground states and tunable hardness across five planting schemes.
-
Volumetric benchmark that scores a device by the largest random Clifford circuit it executes faithfully, using stabilizer simulation to keep verification scalable far past Quantum Volume's ceiling.
-
IBM's speed benchmark: the sustained number of circuit layers per second a quantum system and its classical stack execute while running a batch of parameterized model circuits.
-
Single-number score for error-corrected quantum memories: how many times longer the logical qubit keeps quantum information than the best uncorrected element in the same device, with G = 1 the break-even point.
-
Estimates circuit fidelity by checking how often a device samples the high-probability bitstrings of random quantum circuits, the metric behind Google's quantum supremacy claim.
-
Randomized-measurement protocol that estimates the fidelity between quantum states prepared on two different devices using only classically communicated random unitaries and measurement outcomes
-
Pauli-twirled decay protocol that measures the process fidelity of an entire cycle of parallel gates at once, scaling to whole devices where interleaved benchmarking cannot.
-
Keysight True-Q diagnostic that reconstructs the probabilities of the individual Pauli errors afflicting a cycle of parallel gates, with multiplicative precision
-
Frustrated cluster loops with a tunable coupling scale λ that conceals the planted structure, built to test whether annealer speedups survive against structure-exploiting classical solvers.
-
Suite of Stim-generated syndrome traces for scoring classical QEC decoders on accuracy and latency across surface, color, and bivariate-bicycle codes plus lattice surgery.
-
Deterministic single-qubit protocol that separates coherent from incoherent gate errors with a handful of fixed two-pulse sequences instead of the random circuits of RB.
-
Estimates how close a lab state or gate is to its ideal pure target from a few importance-sampled Pauli measurements, sidestepping the exponential cost of full tomography.
-
Benchmarks random layers of native gates directly instead of compiled multi-qubit Cliffords, measuring gates as they are actually used and scaling well beyond standard RB.
-
The factor by which logical error per QEC cycle drops when code distance rises by two; Λ > 1 certifies below-threshold operation, and Λ anchors Google's error-correction roadmap.
-
Proposed companion to Clifford Volume that scores a device by the largest random free-fermion circuit it executes faithfully, verified scalably through Majorana-mode expectation values.
-
Planted-solution Ising benchmark of the D-Wave 2000Q era: frustrated loops laid over ferromagnetic qubit clusters with tunable ruggedness, scored by time-to-solution against classical solvers.
-
Suite of 160+ algorithm instances compiled into Clifford+T and Pauli-based computation forms to score fault-tolerant resource costs for surface-code and qLDPC architectures.
-
Stim-based suite that scores logical error rates of surface-code Clifford primitives (memory, lattice surgery, transversal Hadamard, and the S gate) under hardware-motivated biased noise.
-
Calibration-free tomography that reconstructs every gate, state preparation, and measurement in a gate set simultaneously and self-consistently, yielding predictive error models rather than a single score.
-
Photonic quantum-advantage benchmark: sample photon-number patterns from squeezed light sent through a large random interferometer, a task whose probabilities are #P-hard matrix hafnians.
-
Whole-device entanglement test that prepares an N-qubit GHZ state and certifies genuine multipartite entanglement whenever the measured fidelity clears 0.5.
-
More than 1.5 million pre-encoded qubit Hamiltonians (spin models, chemistry, and combinatorial optimization) supplying standardized problem instances for application-level quantum benchmarking.
-
Estimates the error of one specific quantum gate by interleaving it between random Cliffords and comparing the resulting decay against a reference randomized benchmarking curve.
-
IonQ's MLPerf-inspired framework scoring whole quantum workloads end to end by solution quality and time-to-solution, wall time from job submission to a result that meets a predefined quality threshold.
-
IBM's scale-friendly quality benchmark: the fidelity of one full layer of simultaneous two-qubit gates across an N-qubit chain, reported per gate as EPLG.
-
Modified randomized benchmarking that tracks population escaping the computational subspace, separating per-gate leakage and seepage rates from ordinary computational error.
-
The headline figure of QEC memory experiments: the probability per round of error correction that the decoded logical qubit suffers a logical flip.
-
Proposed protocols that estimate the infidelity of a prepared logical magic state to multiplicative precision using only fault-tolerantly implementable Clifford operations.
-
Randomized benchmarking lifted to the logical layer: random logical Clifford sequences with QEC cycles interleaved yield an average error per logical gate of an encoded qubit.
-
Google's quick patch-level test that runs a random circuit and its inverse, fits the decay of the return probability, and scores qubit configurations by effective error per cycle.
-
Scores digital and analog quantum processors by the largest system size at which they reproduce quenched transverse-field Ising correlation functions within a chosen error threshold.
-
Scalable RB-style protocol that scores the continuous family of matchgate (XY-type, fermionic-simulation) gates by fitting decays of correlation functions over random matchgate circuits.
-
IBM's randomized-benchmarking suite that quantifies the error a mid-circuit measurement adds to the measured qubit and the dephasing and crosstalk it inflicts on unmeasured spectator qubits.
-
Quantinuum's system-level benchmark that fits the exponential decay of mirrored random circuits' survival probability, with a decay rate that also gauges how coherent the noise is.
-
Scalable RB variant built from mirror circuits (motion-reversal sequences with random Pauli layers) that avoids compiling expensive inversion gates and extends layer-error estimates to many-qubit widths.
-
Library of 70,000+ benchmark circuits, each served at four abstraction levels, so quantum compilers, simulators, and verifiers can be compared on common workloads.
-
Atos/Eviden's application-level metric: the largest MaxCut instance a quantum system can solve effectively, judged by beating random guessing by a set fraction of optimal-solver scaling.
-
PNNL's collection of OpenQASM 2 circuits, spanning chemistry to cryptography at 2 to 433+ qubits, used to evaluate NISQ hardware, compilers, and simulators.
-
DARPA Quantum Benchmarking program's ground-state energy estimation benchmark: classical and quantum solvers are scored on solvability, accuracy, and runtime over a shared library of molecular Hamiltonians.
-
Early hybrid quantum-classical benchmark scoring how well a trained shallow circuit samples the bars-and-stripes distribution, reported as an F1 score of precision and recall.
-
Open-source consortium suite that scores devices on ~16 algorithm workloads, from Bernstein-Vazirani to VQE and MaxCut, via normalized fidelity on width-by-depth volumetric plots.
-
Compiler-stress circuits with provably near-optimal SWAP and depth overheads built in, generalizing QUEKO to benchmark qubit mapping on instances that genuinely need routing
-
Proposed fault-tolerant throughput metric that scores logical operations per second with the classical decoder's accuracy, throughput, and latency folded in.
-
Xanadu's suite for benchmarking quantum machine-learning models against out-of-the-box classical baselines, which won on every task tested.
-
Open library of ten classically hard, practically motivated optimization problem classes (the Intractable Decathlon) with common reporting rules for quantum and classical solvers.
-
TU Delft benchmark suite that runs QAOA and VQE optimization workloads end to end and scores runtime, accuracy, scalability, and capacity.
-
Verification protocol that interleaves a target circuit with randomized Clifford trap circuits to certify its outputs, bounding their variation distance from ideal at chosen confidence
-
Google's second-order OTOC experiment on Willow, framed as a verifiable successor to random-circuit sampling because its echo observable can be re-measured on another quantum computer.
-
Proposed whole-machine benchmark that scores a device on solving linear systems built from random-circuit block-encoded matrices, in the spirit of classical computing's LINPACK.
-
The textbook protocol for fully characterizing a quantum gate: apply it to an informationally complete set of input states, tomograph the outputs, and reconstruct the complete channel.
-
The baseline characterization protocol: reconstructs the full density matrix of a prepared state from repeated measurements in an informationally complete set of bases.
-
Single-number, full-system benchmark that scores a device by the largest random square circuit it can run while still generating heavy outputs reliably.
-
BMW Group's open-source framework for defining, orchestrating, and reproducing application-level quantum benchmarks drawn from industry use cases.
-
Benchmark circuits whose minimal SWAP count on a given device is provably known, scoring layout-synthesis tools by the factor of extra SWAPs they insert over the optimum
-
Synthetic circuits built backward from a device's coupling graph so the optimal mapped depth is known exactly, exposing how far layout-synthesis tools sit from optimal
-
Four-test scalable suite scoring pre-fault-tolerant devices on partial-Clifford random circuits, GHZ-state entanglement, Ising-dynamics simulation reach, and quantum neural network classification accuracy.
-
Estimates average gate error from the exponential decay of survival probability over increasingly long random Clifford sequences, robust to state-preparation and measurement errors.
-
Sandia's scalable whole-processor test that runs self-inverting random circuits across the width × depth plane and maps where a device still returns the right bitstring.
-
The 2008 online database of reversible functions and verified circuit realizations that served as the de facto workload set for quantum compilation benchmarking through the 2010s
-
Heisenberg-limited calibration protocol that pins down individual gate rotation angles and axes from geometrically growing gate repetitions, with provable robustness to SPAM error.
-
Microsoft's proposed FLOPS analogue for fault-tolerant quantum computers: logical qubits times logical clock frequency, qualified by a maximum tolerable logical error rate.
-
Runs randomized benchmarking on several subsystems at once and compares with isolated runs, turning the difference in decay rates into a direct measure of crosstalk and addressability.
-
Reads the decoherence-limited fidelity of a gate straight out of cross-entropy benchmarking data by estimating state purity from the speckle contrast of measured bitstring probabilities.
-
The dual of the QEC memory experiment: instead of preserving a logical observable over time, it scores how well syndrome extraction and the decoder determine a product of stabilizers across space.
-
Scalable, hardware-agnostic suite of eight application-level benchmarks, with a six-dimensional feature vector that profiles how each workload stresses a device.
-
Projected number of physical qubits per logical qubit a QEC code needs, at a stated physical error rate, to reach a one-in-a-trillion logical error rate, extrapolated from simulation rather than measured on hardware.
-
The standard quantum-annealing benchmark: wall-clock time to find the ground state at least once with 99% probability, compared against optimized classical solvers to test for quantum speedup.
-
D-Wave's relaxation of time-to-solution that times solvers to a target energy (originally set by quantiles of a reference annealer's short-run samples) rather than to the exact ground state.
-
Randomized-benchmarking variant that fits the decay of output-state purity to measure how coherent a gate set's noise is, splitting error budgets into miscalibration and decoherence.
-
OpenQASM circuit suite spanning combinational, dynamic, sequential, and variational circuits, built as standard inputs for quantum software tools from equivalence checkers to debuggers
-
Sandia framework that generalizes Quantum Volume from square circuits to the full width × depth plane, mapping the frontier of circuit shapes a device can execute.
No benchmarks match that filter. Try a different term, or add the one you were looking for.
Browse by category
-
Component-level
Characterize individual gates, qubits, and operations in isolation: error rates, coherence, and calibration quality. Home of the randomized benchmarking family.
-
System-level
Exercise a whole processor with structured or random circuits to produce holistic scores that reflect qubit count, fidelity, connectivity, and the compiler together.
-
Application-level
Measure end-to-end performance on programs representative of real workloads, from algorithm subroutines to full application suites.
-
Error correction
Score the error-corrected layer: how fast logical errors fall as codes scale, what logical qubits and fault-tolerant primitives cost, and whether decoders keep up.
-
Software stack
Benchmark the classical software around the QPU (compilers, transpilers, SDKs, and verification tools) on circuit corpora with known baselines or optima.
-
Platform-specific
Benchmarks built for hardware outside the digital gate model: quantum annealers, photonic samplers, and analog quantum simulators.
-
Characterization
Diagnostic protocols (tomography, fidelity estimation, noise learning) that reconstruct what a device actually does: the toolbox that score-style benchmarks build on.