QEncode Benchmark

Benchmark Specification

Suite v4

QEncode uses fixed suite definitions so every run is reproducible and directly comparable across teams. All v4 molecules use the cc-pVDZ basis with chemistry-driven active spaces.

Pipeline: PySCF HF → [CASSCF] → CASCI reference · PennyLane Hamiltonian · Z2 tapering · VQE (COBYLA / L-BFGS-B / Adam)

Suite v4 — Molecule Catalog

16 certified molecules. All use cc-pVDZ basis.

MoleculeActive spaceJW qubitsTaperedEncodingsFlagsStatus
H₂

Hydrogen

[2e, 2o]41
JW
PAR
BK
6 entries
HF

Hydrogen Fluoride

[2e, 2o]41
JW
PAR
BK
6 entries
LiH

Lithium Hydride

[4e, 4o]85
JW
PAR
3 entries
BeH₂

Beryllium Hydride

[4e, 4o]83
JW
PAR
4 entries
H₂O

Water

[4e, 4o]84
JW
PAR
3 entries
NH₃

Ammonia

[4e, 4o]85
JW
PAR
3 entries
H₄

Hydrogen Chain (H₄)

[4e, 4o]85
JW
PAR
4 entries
N₂

Nitrogen

[6e, 6o]128
JW
PAR
CASSCF
3 entries
H₆

Hydrogen Chain (H₆)

[6e, 6o]129
JW
CASSCF
1 entries
H₂CO

Formaldehyde

[4e, 4o]84
JW
PAR
1 entries
C₄H₆

1,3-Butadiene

[4e, 4o]84
JW
PAR
1 entries
(H₂O)₂

Water dimer

[4e, 4o]85
JW
PAR
4 entries
C₄H₄

Cyclobutadiene

[4e, 4o]86
JW
PAR
CASSCF
4 entries
Benzene

Benzene (C₆H₆)

[6e, 6o]129
JW
PAR
CASSCF
2 entries
H₈

Hydrogen Chain (H₈)

[8e, 8o]1613
JW
CASSCF
1 entries
H₁₀

Hydrogen Chain (H₁₀)

[10e, 10o]2018
JW
CASSCF
1 entries

Encoding exclusion notes

BK excluded (LiH, BeH₂, H₂O, NH₃, H₄, N₂, H₆, H₂CO, C₄H₆, (H₂O)₂, C₄H₄, benzene, H₈, H₁₀) — PennyLane 0.45 introduces imaginary artefacts (>7 mHa) in BK tapering for active spaces larger than [2,2]. Only H₂ and HF pass the imaginary-strip check.

PAR/UCCSD excluded (LiH, H₂O, NH₃, N₂, benzene) — UCCSD excitation operators generated in the JW basis are not correctly adapted for Parity tapering in these active spaces. BeH₂ is the exception: D∞h linear symmetry keeps the operator space well-conditioned.

CASSCF required (C₄H₄, N₂, H₆, benzene, H₈, H₁₀) — HF orbitals do not cleanly partition the active space for strongly correlated systems. CASSCF pre-optimises orbitals before the VQE circuit is built.

Qubit Encodings

Jordan-Wigner (JW)

JW

Maps each spin-orbital to one qubit. Locality is preserved along the qubit chain. Supported for all Suite v4 molecules.

Supported: All molecules

Parity (PAR)

PAR

Encodes parity information, enabling 2-qubit reduction. Implemented via OpenFermion bridge. UCCSD operators require care — excluded for LiH, H₂O, NH₃, N₂, benzene.

Supported: All molecules (HEA); select molecules (UCCSD)

Bravyi-Kitaev (BK)

BK

Balances locality and non-locality. Tapering verified clean only for H₂ and HF in PennyLane 0.45. Excluded for all larger molecules due to imaginary artefacts.

Supported: H₂ and HF only

Ansatz Types

UCCSD

Unitary Coupled Cluster Singles and Doubles

Chemically-motivated ansatz. Excitation operators are generated from the active space. High accuracy but deep circuits — parameter count scales with active space size. N₂ UCCSD has 404 parameters.

When to use: Chemically preferred. Required for certified leaderboard entries on strongly correlated molecules.

HEA

Hardware-Efficient Ansatz

Generic parameterised circuit with alternating rotation and entanglement layers. Shallow, hardware-friendly, fast to run. Layer count (reps) is configurable. Sufficient for simple molecules, insufficient for strong multireference systems like N₂.

When to use: Preferred for near-term hardware experiments. Not always sufficient for certification.

ADAPT-VQE

Adaptive Derivative-Assembled Problem-Tailored VQE

Builds the circuit operator by operator, selecting from the UCCSD excitation pool by parameter-shift gradient magnitude at each step. Reaches UCCSD-class accuracy with a small fraction of the parameters, keeping the optimisation tractable where full UCCSD is not.

When to use: Certifies the medium and large molecules — H₂CO, C₄H₆, C₄H₄, H₆, benzene, H₈, H₁₀ — where UCCSD with COBYLA is infeasible.

Reproducibility Guarantees

Every Suite v4 entry is generated by a single script — scripts/generate_entry_v4.py — with a fully pinned environment (requirements-v4.txt). The molecule geometry, basis set, active space, encoding, and ansatz are all predetermined by suite rules.

Each entry is identified by a compound ID such as N2_ccpvdz_JW_UCCSD_v4_casscf_tapered__sha256_82e00cea5a20cd83 and includes a SHA-256 provenance hash and Ed25519 signature for tamper detection. All entries are stored in the public GitHub repository under releases/v4/db/.

Reference energies (HF, MP2, CCSD, CCSD(T), CASCI) are computed by PySCF at the same geometry and basis. VQE gaps are always measured against E_CASCI — never against full-system FCI or a classical approximation.

Read full methodology →View leaderboardGitHub → requirements-v4.txt