Quantum Benchmark Zoo

A community-maintained catalog of protocols for measuring quantum computer performance, from single-gate characterization through whole-processor scores and application suites to error-corrected logical qubits, the software stack, and analog platforms. Each entry covers what the benchmark measures, how it works, key papers, and reference implementations.

All benchmarks

  • Scalable characterization protocol that estimates the Pauli error rates of every gate and measurement on a processor simultaneously from a small set of shallow Clifford circuits

    Flammia (AWS Center for Quantum Computing / Caltech) · 2021 · also known as Averaged circuit eigenvalue sampling, Average circuit eigenvalue sampling

  • IonQ's single-number application metric, the largest width at which a six-algorithm circuit suite clears a 1/e fidelity bar, retired in 2025 in favor of industry-standard metrics.

    IonQ · 2020 · superseded · also known as #AQ, AQ, Algorithmic Qubit

  • Estimates the many-body fidelity of an analog quantum simulator's quench dynamics from about a thousand bitstring samples, doing for Hamiltonian evolution what XEB does for random digital circuits.

    Mark, Choi, Shaw, Endres & Choi (MIT / Caltech / Stanford) · 2022 · also known as Analogue process fidelity, QCMet M10.1, Fidelity estimation via ergodic quantum dynamics, F_d estimator

  • Automated head-to-head comparison of quantum compilers (Qiskit, Tket, Cirq, Quilc, PyZX) on gate counts, depth, circuit cost, and runtime, with auto-generated LaTeX reports

    Kharkov, Ivanova, Mikhantiev & Kotelnikov (Arline / Turation Ltd) · 2020 · historical · also known as arline_benchmarks

  • MetriQs-France benchmark suite scoring whole quantum stacks on optimization, linear systems, many-body simulation, and factoring, aggregated into one user-weighted figure of merit.

    BACQ consortium coordinated by Thales (with Eviden, CEA, CNRS, Teratec & LNE) · 2023 · also known as Application-oriented Benchmarks for Quantum Computing, Benchmarks for Application-Centric Quantum Computing

  • Two-copy benchmark that measures pairs of identical circuit outputs in the Bell basis, yielding fidelity estimates and circuit diagnostics that stay efficient beyond the classically simulable regime.

    Hangleiter & Gullans (QuICS, NIST / University of Maryland) · 2023 · also known as Bell sampling from quantum circuits

  • IBM's open-source suite of 1,000+ timed tests comparing how fast quantum SDKs construct, manipulate, and transpile circuits, plus the quality of the compiled output.

    IBM Quantum · 2024

  • Streamlined, fully scalable RB that drops motion reversal entirely: circuits of i.i.d. random layers act on a random Pauli eigenstate, and each shot returns a simple pass/fail outcome.

    Hines, Proctor & colleagues (Sandia / UC Berkeley) · 2023 · also known as BiRB

  • Uses group character theory to isolate clean exponential decays, extending randomized benchmarking to gate sets that form groups beyond the multi-qubit Cliffords.

    Helsen, Wehner & colleagues (QuTech, Delft) · 2018 · also known as CRB

  • Open-source Python suite that generates Ising, QUBO, and higher-order optimization problems with planted, a-priori-known ground states and tunable hardness across five planting schemes.

    Perera, Katzgraber & colleagues · 2020

  • Volumetric benchmark that scores a device by the largest random Clifford circuit it executes faithfully, using stabilizer simulation to keep verification scalable far past Quantum Volume's ceiling.

    Portik, Kálmán, Monz & Zimborás (EU Quantum Flagship) · 2025 · also known as CLV

  • IBM's speed benchmark: the sustained number of circuit layers per second a quantum system and its classical stack execute while running a batch of parameterized model circuits.

    IBM Quantum (Wack, Paik, Javadi-Abhari & colleagues) · 2021 · also known as Circuit Layer Operations Per Second, CLOPS_v, CLOPS_h, CLOPSv, CLOPSh

  • Single-number score for error-corrected quantum memories: how many times longer the logical qubit keeps quantum information than the best uncorrected element in the same device, with G = 1 the break-even point.

    Sivak, Eickbusch & colleagues (Yale, Devoret group) · 2022 · also known as QEC gain, Lifetime gain, Break-even gain, Beyond-break-even gain

  • Estimates circuit fidelity by checking how often a device samples the high-probability bitstrings of random quantum circuits, the metric behind Google's quantum supremacy claim.

    Google Quantum AI · 2016 · also known as XEB, Linear XEB

  • Randomized-measurement protocol that estimates the fidelity between quantum states prepared on two different devices using only classically communicated random unitaries and measurement outcomes

    Elben, Vermersch, Zoller & colleagues (Innsbruck) · 2019 · also known as Cross-platform fidelity estimation, Cross-platform comparison, Cross-platform fidelity

  • Pauli-twirled decay protocol that measures the process fidelity of an entire cycle of parallel gates at once, scaling to whole devices where interleaved benchmarking cannot.

    Erhard, Wallman & colleagues (Innsbruck & Waterloo) · 2019 · also known as CB

  • Keysight True-Q diagnostic that reconstructs the probabilities of the individual Pauli errors afflicting a cycle of parallel gates, with multiplicative precision

    Quantum Benchmark Inc. (now part of Keysight Technologies), Carignan-Dugas, Emerson, Wallman & colleagues · 2020 · also known as CER, K-body noise reconstruction, KNR

  • Frustrated cluster loops with a tunable coupling scale λ that conceals the planted structure, built to test whether annealer speedups survive against structure-exploiting classical solvers.

    Mandrà (NASA Ames) & Katzgraber (Texas A&M) · 2017 · historical · also known as DCL, deceptive cluster loop problems

  • Suite of Stim-generated syndrome traces for scoring classical QEC decoders on accuracy and latency across surface, color, and bivariate-bicycle codes plus lattice surgery.

    Maurya, Viszlai, Raveendran, Das & Tannu (UW-Madison, UChicago, U. Arizona, UT Austin) · 2025 · also known as Decoder-Bench, DecoderBench

  • Deterministic single-qubit protocol that separates coherent from incoherent gate errors with a handful of fixed two-pulse sequences instead of the random circuits of RB.

    Tripathi, Levenson-Falk, Lidar & colleagues (USC) · 2024 · proposal · also known as DB

  • Estimates how close a lab state or gate is to its ideal pure target from a few importance-sampled Pauli measurements, sidestepping the exponential cost of full tomography.

    Flammia & Liu, da Silva, Landon-Cardinal & Poulin · 2011 · also known as DFE, Monte Carlo quantum process certification

  • Benchmarks random layers of native gates directly instead of compiled multi-qubit Cliffords, measuring gates as they are actually used and scaling well beyond standard RB.

    Proctor, Rudinger & colleagues (Sandia) · 2018 · also known as DRB, Direct RB

  • The factor by which logical error per QEC cycle drops when code distance rises by two; Λ > 1 certifies below-threshold operation, and Λ anchors Google's error-correction roadmap.

    Kelly, Martinis & colleagues (Martinis group, UCSB/Google) · 2014 · also known as Λ, Lambda, lambda factor, exponential suppression factor

  • Proposed companion to Clifford Volume that scores a device by the largest random free-fermion circuit it executes faithfully, verified scalably through Majorana-mode expectation values.

    Portik, Kálmán, Monz & Zimborás (EU Quantum Flagship) · 2025 · proposal · also known as FFV, Free-Fermion Volume

  • Planted-solution Ising benchmark of the D-Wave 2000Q era: frustrated loops laid over ferromagnetic qubit clusters with tunable ruggedness, scored by time-to-solution against classical solvers.

    D-Wave Systems (King, McGeoch & colleagues) · 2017 · historical · also known as FCL, FCL problems, frustrated cluster loop problems

  • Suite of 160+ algorithm instances compiled into Clifford+T and Pauli-based computation forms to score fault-tolerant resource costs for surface-code and qLDPC architectures.

    Harkness et al. (Pacific Northwest National Laboratory & academic collaborators) · 2026 · proposal

  • Stim-based suite that scores logical error rates of surface-code Clifford primitives (memory, lattice surgery, transversal Hadamard, and the S gate) under hardware-motivated biased noise.

    Kan et al. · 2026 · proposal

  • Calibration-free tomography that reconstructs every gate, state preparation, and measurement in a gate set simultaneously and self-consistently, yielding predictive error models rather than a single score.

    Merkel, Gambetta & colleagues (IBM), Blume-Kohout & colleagues (Sandia) · 2012 · also known as GST

  • Photonic quantum-advantage benchmark: sample photon-number patterns from squeezed light sent through a large random interferometer, a task whose probabilities are #P-hard matrix hafnians.

    Hamilton, Kruse, Sansoni, Barkhofen, Silberhorn & Jex · 2016 · also known as GBS

  • Whole-device entanglement test that prepares an N-qubit GHZ state and certifies genuine multipartite entanglement whenever the measured fidelity clears 0.5.

    Community practice; earliest deterministic demonstration by Sackett et al. (NIST) · 2000 · also known as GHZ benchmark, Greenberger–Horne–Zeilinger state fidelity, N-qubit GHZ benchmark

  • More than 1.5 million pre-encoded qubit Hamiltonians (spin models, chemistry, and combinatorial optimization) supplying standardized problem instances for application-level quantum benchmarking.

    Sawaya & colleagues (Intel Labs, LBNL/NERSC, Sandia, NASA Ames, Oxford & others) · 2023 · also known as Hamiltonian Library, HamLib dataset

  • Estimates the error of one specific quantum gate by interleaving it between random Cliffords and comparing the resulting decay against a reference randomized benchmarking curve.

    Magesan, Gambetta & colleagues (IBM / Raytheon BBN) · 2012 · also known as IRB

  • IonQ's MLPerf-inspired framework scoring whole quantum workloads end to end by solution quality and time-to-solution, wall time from job submission to a result that meets a predefined quality threshold.

    IonQ (Aboumrad and colleagues) · 2026 · also known as IonQ Application-Centric Benchmarking Framework, Measuring What Matters, apps-benchmark

  • IBM's scale-friendly quality benchmark: the fidelity of one full layer of simultaneous two-qubit gates across an N-qubit chain, reported per gate as EPLG.

    McKay, Hincks & colleagues (IBM) · 2023 · also known as LF, EPLG, Error per layered gate

  • Modified randomized benchmarking that tracks population escaping the computational subspace, separating per-gate leakage and seepage rates from ordinary computational error.

    Wood & Gambetta (IBM) · 2017 · also known as Leakage RB

  • The headline figure of QEC memory experiments: the probability per round of error correction that the decoded logical qubit suffers a logical flip.

    Google Quantum AI · 2021 · also known as εL, ε_L, logical error per round, logical error per cycle, error per cycle of error correction, per-round logical error rate, LER per round

  • Proposed protocols that estimate the infidelity of a prepared logical magic state to multiplicative precision using only fault-tolerantly implementable Clifford operations.

    Lee, Yuan, Chen, Tsubouchi & Jiang (University of Chicago / University of Tokyo) · 2025 · proposal · also known as Efficient benchmarking of logical magic state, Logical magic state infidelity benchmarking

  • Randomized benchmarking lifted to the logical layer: random logical Clifford sequences with QEC cycles interleaved yield an average error per logical gate of an encoded qubit.

    Combes, Granade, Ferrie & Flammia · 2017 · proposal · also known as LRB, Logical RB

  • Google's quick patch-level test that runs a random circuit and its inverse, fits the decay of the return probability, and scores qubit configurations by effective error per cycle.

    Google Quantum AI · 2021 · also known as Loschmidt echoes, loschmidt.tilted_square_lattice, Qubit picking with Loschmidt echoes

  • Scores digital and analog quantum processors by the largest system size at which they reproduce quenched transverse-field Ising correlation functions within a chosen error threshold.

    Erbin, Burdeau, Bertrand, Ayral & Misguich (Eviden Quantum Lab & IPhT, Université Paris-Saclay) · 2026 · proposal · also known as MBQS

  • Scalable RB-style protocol that scores the continuous family of matchgate (XY-type, fermionic-simulation) gates by fitting decays of correlation functions over random matchgate circuits.

    Helsen, Nezami, Reagor & Walter · 2020 · also known as Matchgate RB, Matchgate randomized benchmarking

  • IBM's randomized-benchmarking suite that quantifies the error a mid-circuit measurement adds to the measured qubit and the dephasing and crosstalk it inflicts on unmeasured spectator qubits.

    Govia, Jurcevic & colleagues (IBM Quantum) · 2022 · also known as Mid-circuit measurement randomized benchmarking, Mid-circuit measurement RB, MCM-RB, mcm-rb suite, MCM benchmarking

  • Quantinuum's system-level benchmark that fits the exponential decay of mirrored random circuits' survival probability, with a decay rate that also gauges how coherent the noise is.

    Mayer, Hall & colleagues (Honeywell Quantum Solutions, now Quantinuum) · 2021 · also known as MB, Quantinuum mirror benchmarking, Honeywell mirror benchmarking

  • Scalable RB variant built from mirror circuits (motion-reversal sequences with random Pauli layers) that avoids compiling expensive inversion gates and extends layer-error estimates to many-qubit widths.

    Proctor, Seritan & colleagues (Sandia) · 2021 · also known as MRB, Mirror RB

  • Library of 70,000+ benchmark circuits, each served at four abstraction levels, so quantum compilers, simulators, and verifiers can be compared on common workloads.

    Quetschlich, Burgholzer & Wille (TU Munich, Chair for Design Automation) · 2022

  • Atos/Eviden's application-level metric: the largest MaxCut instance a quantum system can solve effectively, judged by beating random guessing by a set fraction of optimal-solver scaling.

    Atos (Martiel, Ayral & Allouche) · 2020 · also known as Atos Q-score, Eviden Q-score, Q-Score, Q-score MaxCut, Q-score Max-Clique, Qs

  • PNNL's collection of OpenQASM 2 circuits, spanning chemistry to cryptography at 2 to 433+ qubits, used to evaluate NISQ hardware, compilers, and simulators.

    Li, Stein, Krishnamoorthy & Ang (Pacific Northwest National Laboratory) · 2020 · also known as QASMBench suite, PNNL QASMBench

  • DARPA Quantum Benchmarking program's ground-state energy estimation benchmark: classical and quantum solvers are scored on solvability, accuracy, and runtime over a shared library of molecular Hamiltonians.

    DARPA Quantum Benchmarking program (Bellonzi et al.: L3Harris, Zapata AI, University of Toronto, HRL & others) · 2024 · also known as QB Ground State Energy Estimation Benchmark, QB-GSEE-Benchmark, QB-GSEE

  • Early hybrid quantum-classical benchmark scoring how well a trained shallow circuit samples the bars-and-stripes distribution, reported as an F1 score of precision and recall.

    Benedetti, Perdomo-Ortiz & colleagues · 2018 · historical · also known as qBAS, qBAS(n,m), qBAS22

  • Open-source consortium suite that scores devices on ~16 algorithm workloads, from Bernstein-Vazirani to VQE and MaxCut, via normalized fidelity on width-by-depth volumetric plots.

    QED-C (Quantum Economic Development Consortium) · 2021 · also known as QED-C Application-Oriented Performance Benchmarks, Application-Oriented Performance Benchmarks, QC-App-Oriented-Benchmarks, QED-C benchmarks

  • Compiler-stress circuits with provably near-optimal SWAP and depth overheads built in, generalizing QUEKO to benchmark qubit mapping on instances that genuinely need routing

    Li, Zhou & Feng (UTS, Nanjing Tech & Tsinghua) · 2023 · also known as QUEKNO, Qubit Mapping Benchmark with Known Near-Optimality

  • Proposed fault-tolerant throughput metric that scores logical operations per second with the classical decoder's accuracy, throughput, and latency folded in.

    Kong, Zhang & Chen (Zhongguancun Laboratory & Tsinghua University) · 2025 · proposal · also known as Quantum Logical Operations Per Second, logical operations per second

  • Xanadu's suite for benchmarking quantum machine-learning models against out-of-the-box classical baselines, which won on every task tested.

    Bowles, Ahmed & Schuld (Xanadu) · 2024 · also known as Better than classical?, Xanadu QML benchmarks

  • Open library of ten classically hard, practically motivated optimization problem classes (the Intractable Decathlon) with common reporting rules for quantum and classical solvers.

    Quantum Optimization Working Group (Zuse Institute Berlin, IBM Quantum & partners) · 2025 · also known as Quantum Optimization Benchmarking Library, Quantum Optimization Benchmark Library, Intractable Decathlon

  • TU Delft benchmark suite that runs QAOA and VQE optimization workloads end to end and scores runtime, accuracy, scalability, and capacity.

    Mesman, Al-Ars & Möller (TU Delft) · 2021 · historical · also known as QPack Scores

  • Verification protocol that interleaves a target circuit with randomized Clifford trap circuits to certify its outputs, bounding their variation distance from ideal at chosen confidence

    Ferracin, Kapourniotis & Datta (University of Warwick) · 2018 · also known as Accreditation protocol, AP

  • Google's second-order OTOC experiment on Willow, framed as a verifiable successor to random-circuit sampling because its echo observable can be re-measured on another quantum computer.

    Google Quantum AI and collaborators (Abanin et al.) · 2025 · also known as OTOC(2), OTOC^(2), Second-order out-of-time-order correlator benchmark

  • Proposed whole-machine benchmark that scores a device on solving linear systems built from random-circuit block-encoded matrices, in the spirit of classical computing's LINPACK.

    Dong & Lin (UC Berkeley) · 2020 · proposal · also known as Quantum LINPACK benchmark, RACBEM

  • The textbook protocol for fully characterizing a quantum gate: apply it to an informationally complete set of input states, tomograph the outputs, and reconstruct the complete channel.

    Chuang & Nielsen, Poyatos, Cirac & Zoller · 1996 · also known as QPT, Process tomography, Standard quantum process tomography, SQPT

  • The baseline characterization protocol: reconstructs the full density matrix of a prepared state from repeated measurements in an informationally complete set of bases.

    Vogel & Risken (theory), Smithey, Beck, Raymer & Faridani (first experiment), James, Kwiat, Munro & White (standard qubit recipe) · 1989 · also known as QST, State tomography, Density-matrix reconstruction, Quantum state estimation

  • Single-number, full-system benchmark that scores a device by the largest random square circuit it can run while still generating heavy outputs reliably.

    IBM · 2018 · also known as QV

  • BMW Group's open-source framework for defining, orchestrating, and reproducing application-level quantum benchmarks drawn from industry use cases.

    Finžgar, Ross, Klepsch & Luckow (BMW Group) · 2022 · also known as QUantum computing Application benchmaRK, QUARK framework

  • Benchmark circuits whose minimal SWAP count on a given device is provably known, scoring layout-synthesis tools by the factor of extra SWAPs they insert over the optimum

    Ping, Lin, Tan & Cong (UCLA) · 2025 · also known as QUantum Benchmark wIth Known-Optimal SWAP Counts

  • Synthetic circuits built backward from a device's coupling graph so the optimal mapped depth is known exactly, exposing how far layout-synthesis tools sit from optimal

    Tan & Cong (UCLA VAST lab) · 2020 · also known as QUEKO benchmarks, QUantum Mapping Examples with Known Optimal

  • Four-test scalable suite scoring pre-fault-tolerant devices on partial-Clifford random circuits, GHZ-state entanglement, Ising-dynamics simulation reach, and quantum neural network classification accuracy.

    Aguirre, Peña & Sanz (University of the Basque Country) · 2025 · proposal · also known as QuSquare Benchmark Suite

  • Estimates average gate error from the exponential decay of survival probability over increasingly long random Clifford sequences, robust to state-preparation and measurement errors.

    Emerson, Alicki & Życzkowski, Knill et al. (NIST), Magesan, Gambetta & Emerson · 2005 · also known as RB, Clifford RB

  • Sandia's scalable whole-processor test that runs self-inverting random circuits across the width × depth plane and maps where a device still returns the right bitstring.

    Proctor, Rudinger, Young, Nielsen & Blume-Kohout (Sandia) · 2020 · also known as Mirror circuit benchmarks, Mirror circuits, Mirrored circuits average polarization, Periodic mirror circuits

  • The 2008 online database of reversible functions and verified circuit realizations that served as the de facto workload set for quantum compilation benchmarking through the 2010s

    Wille, Große, Teuber, Dueck & Drechsler (U. Bremen & U. New Brunswick) · 2008 · superseded · also known as RevLib benchmarks

  • Heisenberg-limited calibration protocol that pins down individual gate rotation angles and axes from geometrically growing gate repetitions, with provable robustness to SPAM error.

    Kimmel, Low & Yoder · 2015 · also known as RPE

  • Microsoft's proposed FLOPS analogue for fault-tolerant quantum computers: logical qubits times logical clock frequency, qualified by a maximum tolerable logical error rate.

    Microsoft Azure Quantum (Krysta Svore) · 2023 · proposal · also known as Reliable Quantum Operations Per Second, reliable QOPS, RQOPS

  • Runs randomized benchmarking on several subsystems at once and compares with isolated runs, turning the difference in decay rates into a direct measure of crosstalk and addressability.

    Gambetta, Córcoles & colleagues (IBM / Raytheon BBN) · 2012 · also known as SRB

  • Reads the decoherence-limited fidelity of a gate straight out of cross-entropy benchmarking data by estimating state purity from the speckle contrast of measured bitstring probabilities.

    Arute et al. (Google AI Quantum) · 2019 · also known as SPB, Purity XEB, XEB purity

  • The dual of the QEC memory experiment: instead of preserving a logical observable over time, it scores how well syndrome extraction and the decoder determine a product of stabilizers across space.

    Craig Gidney (Google Quantum AI) · 2022

  • Scalable, hardware-agnostic suite of eight application-level benchmarks, with a six-dimensional feature vector that profiles how each workload stresses a device.

    Super.tech (now Infleqtion), University of Chicago (EPiQC) · 2022

  • Projected number of physical qubits per logical qubit a QEC code needs, at a stated physical error rate, to reach a one-in-a-trillion logical error rate, extrapolated from simulation rather than measured on hardware.

    Gidney, Newman, Fowler & Broughton (Google Quantum AI) · 2021 · also known as Teraquop qubit count, Footprint of a teraquop

  • The standard quantum-annealing benchmark: wall-clock time to find the ground state at least once with 99% probability, compared against optimized classical solvers to test for quantum speedup.

    Rønnow, Lidar, Troyer & colleagues · 2014 · also known as TTS, TTS99, time-to-99%-success-probability

  • D-Wave's relaxation of time-to-solution that times solvers to a target energy (originally set by quantiles of a reference annealer's short-run samples) rather than to the exact ground state.

    King, Yarkoni, Nevisi, Hilton & McGeoch (D-Wave) · 2015 · also known as TTT, TTT_total, TTT_anneal

  • Randomized-benchmarking variant that fits the decay of output-state purity to measure how coherent a gate set's noise is, splitting error budgets into miscalibration and decoherence.

    Wallman, Granade, Harper & Flammia · 2015 · also known as Unitarity RB, Purity Benchmarking, Purity RB, Purity Randomized Benchmarking, XRB, Extended Randomized Benchmarking

  • OpenQASM circuit suite spanning combinational, dynamic, sequential, and variational circuits, built as standard inputs for quantum software tools from equivalence checkers to debuggers

    Chen, Ying & colleagues (Veri-Q project) · 2022 · also known as Veri-Q Benchmark

  • Sandia framework that generalizes Quantum Volume from square circuits to the full width × depth plane, mapping the frontier of circuit shapes a device can execute.

    Blume-Kohout & Young (Sandia) · 2019

Browse by category

  • Component-level

    Characterize individual gates, qubits, and operations in isolation: error rates, coherence, and calibration quality. Home of the randomized benchmarking family.

  • System-level

    Exercise a whole processor with structured or random circuits to produce holistic scores that reflect qubit count, fidelity, connectivity, and the compiler together.

  • Application-level

    Measure end-to-end performance on programs representative of real workloads, from algorithm subroutines to full application suites.

  • Error correction

    Score the error-corrected layer: how fast logical errors fall as codes scale, what logical qubits and fault-tolerant primitives cost, and whether decoders keep up.

  • Software stack

    Benchmark the classical software around the QPU (compilers, transpilers, SDKs, and verification tools) on circuit corpora with known baselines or optima.

  • Platform-specific

    Benchmarks built for hardware outside the digital gate model: quantum annealers, photonic samplers, and analog quantum simulators.

  • Characterization

    Diagnostic protocols (tomography, fidelity estimation, noise learning) that reconstruct what a device actually does: the toolbox that score-style benchmarks build on.