Skip to content

· VerifyCore Labs

Treat scientific disagreement as a queue of questions before it becomes a verdict

Turn conflicting scientific evidence into reproducible screening decisions.

Who did this. The lab’s AI agents did the research and engineering. Nick Harris, founder. CTO of VivaMed BioPharma; co-founder of MedSim.ai, FastRead.io and Formulai. The lab’s track record.

A scientific-data pipeline can make a conclusion look stronger by counting several agreeing methods. That count is meaningful only if the methods answer the same question and do not inherit the same blind spot. Different software labels are not enough to establish independent evidence.

For an R&D team, the practical task is to turn disagreement and apparent agreement into reviewable decisions: which candidate advances, which experiment resolves uncertainty, and which claim must remain unknown. The supporting record needs to preserve corrections as carefully as positive results.

A public correction worth learning from

Lattice Graph's lithium-hafnate article describes calculations initially interpreted as instability, followed by literature showing synthesis history. The post says the DFT and machine-learned results were not judged on identical frequency sets; under its stricter criterion the DFT verdict is unknown. It also states that saved forces are missing, preventing exact repetition from the stored result.

The public source record names inputs and literature references. These are the lab's calculations and review, not peer-reviewed validation of a mechanism explaining the discrepancy. The appropriate lesson is a change in workflow, not a claim that computation or experiment is categorically superior.

Write the question before comparing answers

“Is this material stable?” can mean several things. A calculation may address a particular structure under specified assumptions. A synthesis paper may show a material with the same composition was made under particular conditions. Neither answer automatically resolves every structure, temperature, defect concentration, or intended application.

A review record should therefore begin with a complete question: structure identifier, property, conditions, method, and decision threshold. When importing a result, record which parts match that question. A composition-only literature match is a lead for reading, not confirmation of an identical phase.

Use matching to locate evidence

This keeps an automated pipeline from turning a convenient join key into a scientific conclusion. Formula matching is useful for finding candidate sources; interpreting those sources remains a separate step.

An example of agreement that needs unpacking

Imagine a screening table with three models predicting a candidate will fail and one experimental report suggesting a usable material. A majority vote would reject it. A better first review asks whether the models share training data, whether they used the same structure, and whether the reported experiment measured the same property.

Create an evidence matrix with one row per source. Columns should distinguish direct measurement, calculated value, literature interpretation, missing inputs, and unresolved assumptions. Then write the decision being proposed and the evidence that would reverse it.

List shared assumptions as questions

If all calculations inherit the same starting structure, a structure mismatch is a candidate explanation. If all exclude finite-temperature effects, that exclusion is another question to test. Neither becomes the explanation merely because it sounds plausible.

The cheapest useful next step might be reading a full paper, inspecting a structure file, checking a convergence parameter, or rerunning a small calculation with preserved inputs. Expensive computation should follow the unresolved question rather than substitute for defining it.

Negative findings need scope and expiration

A negative result can save a team time, but only if it describes what failed. “Material X failed” is too broad when the observed failure concerns one structure, processing condition, assay, or simulation setup.

Retain the attempted method, conditions, acceptance criterion, and outcome. Distinguish a completed negative experiment from a failed pipeline execution. A calculation that did not converge does not establish the candidate lacks the desired property.

Preserve revisions to negative findings

Give each finding a review status and a reason for changing it. New evidence may confirm, narrow, or rebut the result. Preserve the earlier decision so someone can understand why it was made with the evidence available then.

Lattice Graph's public description of its negative-findings record uses grades and later-status changes, but discloses that grade assignment and sample entries are not public. A buyer should ask to inspect those details before relying on the grades in a screening decision. A count of entries is not an estimate of their predictive value.

Preserve the evidence needed for a rerun

A summary value is often enough to populate a dashboard and insufficient to repeat the work. Decide which inputs are necessary before launching the calculation: source structure, transformations, parameter files, raw outputs, software revision, environment, and interpretation code.

Store the decision rule beside the run. If the rule later changes, recompute the verdict from preserved evidence where possible. If the inputs do not permit that, label the result unavailable under the new rule rather than retrospectively treating it as checked.

Ask another engineer to replay the decision

A useful recovery exercise asks a second engineer to rebuild a small set of decisions using only the retained review package. Record missing inputs and divergent interpretations. That is more informative than checking that every row has a nonempty provenance field.

Audit the decision boundary

Before a candidate enters a costly experimental stage, ask:

  • Do all compared sources concern the same structure, property, and conditions?
  • Which methods share data, assumptions, or preprocessing?
  • Is absence from a database being mistaken for novelty?
  • Are failed executions being mistaken for negative scientific evidence?
  • Can the verdict be reconstructed from retained inputs and an explicit rule?
  • What small experiment or source review would change the decision?
  • Are corrections visible to downstream users who received the original result?

These questions create an auditable boundary between data processing and scientific judgment. They do not require a pipeline to be infallible. They require it to reveal where certainty ends and a decision begins.

A concrete evaluation offer

For a materials or scientific-software team, a practical first consulting engagement is an audit of one screening cohort: review source alignment, trace decision rules, test replayability, and identify the smallest checks that could alter high-value decisions. Proposed deliverables include an evidence matrix, a correction log, a reproducibility assessment, and a prioritized experiment queue.

Discuss a scoped evaluation

Contact nick@latticegraph.com to discuss that scope or evaluation of Lattice Graph research materials. VerifyCore uses AI agents for research and engineering. Public calculations do not establish clinical utility, patent novelty, customer adoption, or commercial value. Those questions need their own evidence and qualified review; licensing or acquisition discussions also require rights and dependency diligence.

All postsAll resultsDiscuss an evaluation

How we show numbers

A number you can click opens the file it comes from. How each result is checked

  • We never show a number before its file has loaded.
  • A question we have not checked yet is marked as unchecked.
  • A check that found nothing says so.
  • A file with no value for a question says so.
  • A number whose file is missing or has changed is not shown.
  • Two files that disagree about what a number describes are both flagged.
  • A number from too few samples shows its sample size.
  • Two files that give different values are both shown.
  • A file we cannot publish is listed by its fingerprint only.
  • Where a file records when it was measured, the date is in that file.
  • A question that does not apply to a page is left off it.
  • A measurement whose program failed is shown as failed.