No review leakage
eLife review sub-articles, history, links and family identifiers were removed and audited again at execution. Zero leak checks failed.
We froze and blind-ran 100 eLife papers and 100 medRxiv papers. The evaluation infrastructure held. The current product is not yet a scientific reviewer.
release 2026-08-17 · predictions locked before pilot labels · sealed outcomes unopened
Every non-sealed eLife and medRxiv paper was held as incomplete. That avoids a dangerous false green, but it does not prioritize papers or identify the decisive scientific problem.
not_implemented. A post-release candidate layer is
now live, but both 20-paper pilot queues still contain zero adjudicated papers. Two domain
reviewers and an adjudicator must create the gold standard before precision or recall exists.
The new Review Lab extracts exact claims and issue candidates, checks
methods and statistical reporting, retrieves candidate evidence, and compares revisions.
Every heuristic remains capped at REVIEW.
Each source uses an exact, cohort-stratified 20 pilot / 50 development / 15 validation / 15 sealed-test allocation. Prediction files were SHA-256 locked before pilot review or annotation scaffolds were generated.
| Corpus | Selection | Private bundle | Pilot work |
|---|---|---|---|
| eLife | 50 multi-version · 25 incomplete/inadequate · 25 solid+ | 162/163 versions; all 100 v1 inputs | 133 review documents · 749 weak candidates |
| medRxiv | 25 v1 · 25 revised · 25 published · 15 trial · 10 withdrawn | 146/146 versions across 47 categories | 20 papers · 10 version pairs · no prefilled claims |
eLife review sub-articles, history, links and family identifiers were removed and audited again at execution. Zero leak checks failed.
All 308 executed versions produced independently revalidated source, report, decision and receipt hashes.
The medRxiv selector used requester-pays S3 byte ranges to fetch only matching MECA XML members. Full text remains private and licence-stamped.
The existing desk screen is a bibliographic and integrity preflight. It frequently detects unidentifiable references and uncheckable high-risk prose, but it does not reason over study design, methods, statistics, controls or claim support.
not_implementednot_implementednot_implementednot_implementedExact quote, locator, typed proposition, uncertainty and high-recall candidate generation.
Design, statistical unit, replication, controls, multiplicity, sensitivity and claim/design mismatch.
Align review issues to claims, compare versions semantically and require evidence for resolution.
Stages 01–03 plus revision comparison and dual-review/adjudication UI are now implemented. The next gate is to score expert pilot and development labels, freeze thresholds on validation, and open the sealed set once.
Biosingularity should be the evidence-risk control layer between literature or analysis and a consequential decision: what is claimed, what exact record supports it, what changed, and what is unsafe to assert.
Translational R&D, scientific affairs, evidence strategy, research integrity and R&D data/AI governance.
Pre-submission checks, due diligence, editorial triage, systematic review and portfolio evidence monitoring.
Exact claim receipts, fail-closed decisions, version-aware reverification and transparent peer-review evaluation.
No manuscript or review text is published. The repository contains the methods, schemas, aggregate report, corpus/split ids and prediction hashes.
Full technical findings →
Open the no-human proxy validation dashboard →
Expert review protocol →
medRxiv TDM policy → ·
eLife terms →