Skip to content

· VerifyCore Labs

When two outside solvers disagree, keep the disagreement in the benchmark

A review package that preserves reference-solver disagreement and cohort limits.

Who did this. The lab’s AI agents did the research and engineering. Nick Harris, founder. CTO of VivaMed BioPharma; co-founder of MedSim.ai, FastRead.io and Formulai. The lab’s track record.

A numerical model looks more convincing when it is compared with software its author did not write. That comparison is useful, but it still requires judgment. Outside software can be configured incorrectly, solve a different boundary problem, or disagree with another outside program. Independence of who wrote the code is one part of validation; agreement on the measured quantity is another.

For a package-extraction buyer, the most valuable benchmark may be the one that makes disagreement easiest to inspect. A single improvement ratio can hide more than it explains if its geometries, reference quantities, and exclusions are missing.

What ChipletOS makes public

The ChipletOS comparison page reports that its fast coupling model was closer to FastCap and Palace than a pairwise baseline the lab defined, on sampled capacitance populations. Palace covered one slice. The lab ran the outside programs itself; no outside firm audited the result.

The published result receipt preserves several populations rather than replacing them with one headline. It names checking commands and a negative control. Its scope excludes an independent inductance anchor. The site also discloses unresolved differences in absolute accuracy between the two outside solvers. These are simulation comparisons, not silicon measurements or evidence of outperforming a commercial extraction suite.

First agree on what is being compared

Consider a hypothetical via array. One program reports a capacitance matrix, another exports a reduced port quantity, and the fast model predicts a coupling operator. Those objects may be related, but the relationship must be explicit before computing an error.

A reviewer should ask how conductor ordering, reference potential, sign conventions, units, and eliminated conductors are handled. A permutation can make the right matrix appear wrong. A unit conversion error can make a large improvement look plausible if both candidate and baseline share it.

Check the output adapters first

Use a tiny analytic or symmetric case to test the adapters before testing the interesting geometry. Record the raw solver output alongside the normalized representation. Keep the normalization script small enough to review independently of the fast model. This separates adapter correctness from model accuracy.

These recommendations are an evaluation design. They do not assert that the lab has already completed every adapter check described here.

A baseline describes a question

A pairwise sum asks how a particular simplification compares with a model that includes more interactions. That is a useful research question. It does not automatically describe what commercial solvers do, nor establish competitive advantage over a buyer's existing workflow.

Name the baseline in the claim itself. If the question is “does this model improve on our pairwise approximation for these layouts?”, report that. If a buyer wants to know whether it can replace an extraction stage, the evaluation must include the buyer's stage, permitted error, runtime budget, and relevant geometry distribution.

Give the baseline a useful purpose

A baseline should also have a chance to succeed. If it is intentionally crude, explain why that crudeness represents an actual proposed use. Otherwise the comparison may identify a weakness nobody intends to ship.

Keep cohorts and exclusions separate

Suppose one campaign contains easy layouts and another emphasizes strong coupling. An average error from the first and an improvement ratio from the second should not be combined into a single population claim. Different cohorts can answer different questions without one superseding the other.

Before running an evaluation, define the geometry families, parameter bounds, reference tools, and inclusion rules. Assign each geometry a stable identifier. An excluded point should retain its identifier and reason: convergence failure, invalid mesh, unsupported configuration, or a rule decided in advance.

Retain failed reference solves

A failing reference solve is not a favorable candidate outcome. Nor should a finite subset be quietly promoted into a claim about a full geometry family. Report the number attempted, the number graded, and the number excluded for each reference tool.

This also prevents accidental selection after observing errors. If a coupling-heavy slice is the intended pilot, call it that and preserve it. A later broad campaign needs its own manifest and acceptance threshold.

Study disagreement as an output

If two reference solvers disagree, do not pick whichever gives the prettier headline. Break the difference into testable hypotheses: mesh resolution, computational-domain extent, conductor representation, boundary conditions, and conversion to the comparison quantity.

Change one factor at a time on a small fixed cohort. Preserve raw outputs, normalized outputs, software versions, convergence criteria, and runtime. Label explanations that have not been tested as hypotheses. An attractive physical story is not a measured decomposition.

Show the discrepancy for each case

For each geometry, show candidate-to-reference error against both programs and reference-to-reference discrepancy. The last column tells a buyer whether the candidate's apparent error is small compared with disagreement among graders. It does not excuse error; it makes the uncertainty visible.

If disagreement remains, the decision may be to narrow the supported domain or defer a claim. That is a productive outcome when it prevents a model from being used outside its demonstrated range.

A procurement-ready review package

A compact review package should include:

  • A geometry manifest with units, material assumptions, boundaries, and inclusion rules.
  • Raw reference outputs and the reviewed normalization path.
  • Per-geometry candidate, baseline, and reference discrepancies.
  • Convergence studies on representative and difficult cases.
  • Failed runs and excluded points with reasons.
  • Runtime and memory measurements taken under comparable conditions, if those are being claimed.
  • A statement of quantities that remain internally graded or untested.

A file fingerprint can establish that a served artifact matches a published commitment. It cannot establish that the physical problem was configured correctly or rerun a solver. Those require separate review and execution.

A concrete evaluation offer

For an EDA or signal-integrity team, a useful first engagement is a scoped benchmark audit or integration study using a buyer-selected public or shareable geometry set. The proposed deliverables are a frozen cohort, reference adapters, per-case discrepancies, failure records, and a go/no-go report against thresholds selected before execution.

Contact nick@latticegraph.com to discuss that consulting scope or evaluation of ChipletOS research materials. VerifyCore's research and engineering use AI agents. Any licensing or acquisition discussion would separately assess rights, dependencies, reproducibility, and the gap between a research model and the buyer's production workflow. This article claims no measured commercial saving or customer deployment.

All postsAll resultsDiscuss an evaluation

How we show numbers

A number you can click opens the file it comes from. How each result is checked

  • We never show a number before its file has loaded.
  • A question we have not checked yet is marked as unchecked.
  • A check that found nothing says so.
  • A file with no value for a question says so.
  • A number whose file is missing or has changed is not shown.
  • Two files that disagree about what a number describes are both flagged.
  • A number from too few samples shows its sample size.
  • Two files that give different values are both shown.
  • A file we cannot publish is listed by its fingerprint only.
  • Where a file records when it was measured, the date is in that file.
  • A question that does not apply to a page is left off it.
  • A measurement whose program failed is shown as failed.