Where Browsers Differ · stage 06
Testing what you cannot see
Conformance suites exist because agreement has to be demonstrated, not assumed.
Stage 06 of six/Where Browsers Differ, piece 3 of 3/Full piece
When a specification is not enough
A specification is a document, not a guarantee. It can be read, interpreted, and implemented by several independent teams, all of whom will make different assumptions about the cases it left ambiguous. The result is browsers that agree on everything the spec made explicit and diverge quietly everywhere else. Conformance suites exist to make that divergence visible — to turn a prose document into a battery of machine-verifiable claims.
The idea is simple: if a behaviour can be specified, it can also be expressed as a test that either passes or fails. Assemble enough of those tests, run them against every implementation, and you get a map of where the spec ends and interpretation begins. The W3C has maintained this discipline for decades. Its test suite infrastructure — now housed largely in the web-platform-tests repository — contains hundreds of thousands of individual tests covering HTML, CSS, APIs, and the lower-level platform behaviours that web pages depend on. Every major browser engine contributes tests and runs the full suite against its own builds continuously.
What a conformance test actually does
A conformance test does one thing: it constructs a situation the spec makes a claim about and checks whether the implementation matches. In the CSS world, many of these tests are reference tests — a page rendered by the implementation under test is pixel-compared against a reference rendering known to be correct. If the bitmaps match, the test passes. If they diverge even by a few pixels, something is wrong with at least one of them. This approach is blunt but honest; it catches the cases where two engines produce visually different output without either generating a JavaScript error or throwing an obvious exception.
Other tests are scripted: JavaScript that exercises an API, checks return values, and reports pass or fail through a standard harness called testharness.js. These are particularly useful for behaviour that has no visual output — event ordering, attribute reflection, parsing edge cases. A CSS @layer ordering rule, a fetch() that should reject under a particular security policy, a custom element lifecycle callback that must fire in a specific sequence: all of these can be expressed as assertions a machine can evaluate in milliseconds.
The harder category is behaviour that is deliberately underspecified. Specifications sometimes grant implementations latitude — an animation easing function can produce values within a tolerance, a heuristic can be implementation-defined, pixel-snapping may vary by platform. Tests in these zones cannot demand exact agreement; they can only constrain the range of acceptable answers. Writing those constraints well is its own discipline, and getting them wrong produces tests that pass everywhere while masking real divergence, or tests that fail in Firefox but not Chrome for reasons that turn out to be font rasterisation rather than CSS logic.
The gap between passing and correct
A conformance suite is a sample, not a proof. Passing every test in the suite demonstrates that an implementation handled every case the test authors thought to express — which is a different claim from handling every case the spec intended. Gaps in test coverage are gaps in demonstrated conformance, and those gaps are exactly where the same spec produces two results in practice.
This is why the web-platform-tests project is treated as a living document rather than a fixed artefact. When a browser bug is reported, the fix typically arrives alongside a new test that encodes the correct behaviour, so the bug cannot silently regress and other engines can verify they never had the same flaw. The suite grows toward completeness the way proofs grow toward certainty — never quite arriving, but narrowing the unknown with each iteration.
The practical consequence for anyone building for the web is that passing the suite is a floor, not a ceiling. An engine that passes all current tests is still making undocumented choices wherever the tests are thin. Those choices show up in production: a stacking context that works in three browsers and breaks in one, a grid placement edge case that the spec resolved but the test suite has not yet encoded. Understanding that conformance testing is demonstrative rather than exhaustive is what separates treating a browser bug as a surprise from treating it as an expected feature of a system in which agreement has to be earned, continuously, one test at a time.
Tests in these zones cannot demand exact agreement; they can only constrain the range of acceptable answers.
The harder category is behaviour that is deliberately underspecified.
The stage in a room: conformance suites exist because agreement has to be demonstrated, not assumed.