biosingularity

What the papers told us

We froze and blind-ran 100 eLife papers and 100 medRxiv papers. The evaluation infrastructure held. The current product is not yet a scientific reviewer.

release 2026-08-17 · predictions locked before pilot labels · sealed outcomes unopened

200frozen paper families
308private version artifacts executed
0execution failures or invalid receipts
170 / 170non-sealed papers classified incomplete

Safe abstention is working. Scientific review is not.

Every non-sealed eLife and medRxiv paper was held as incomplete. That avoids a dangerous false green, but it does not prioritize papers or identify the decisive scientific problem.

No scientific accuracy claim is being made. At prediction lock, claim and scientific-issue review were not_implemented. A post-release candidate layer is now live, but both 20-paper pilot queues still contain zero adjudicated papers. Two domain reviewers and an adjudicator must create the gold standard before precision or recall exists.

The four missing stages are now runnable

The new Review Lab extracts exact claims and issue candidates, checks methods and statistical reporting, retrieves candidate evidence, and compares revisions. Every heuristic remains capped at REVIEW.

40unsealed pilot papers rerun
5,101 / 5,101exact claim source locators
181methods review prompts
0heuristic ALLOW outcomes

Open the Review Lab →
Read the candidate-pilot findings →

Frozen before we looked

Each source uses an exact, cohort-stratified 20 pilot / 50 development / 15 validation / 15 sealed-test allocation. Prediction files were SHA-256 locked before pilot review or annotation scaffolds were generated.

CorpusSelectionPrivate bundlePilot work
eLife50 multi-version · 25 incomplete/inadequate · 25 solid+162/163 versions; all 100 v1 inputs133 review documents · 749 weak candidates
medRxiv25 v1 · 25 revised · 25 published · 15 trial · 10 withdrawn146/146 versions across 47 categories20 papers · 10 version pairs · no prefilled claims

What worked

Blindness

No review leakage

eLife review sub-articles, history, links and family identifiers were removed and audited again at execution. Zero leak checks failed.

Provenance

Receipts stayed valid

All 308 executed versions produced independently revalidated source, report, decision and receipt hashes.

Acquisition

Bulk TDM without rehosting

The medRxiv selector used requester-pays S3 byte ranges to fetch only matching MECA XML members. Full text remains private and licence-stamped.

What the frozen baseline could not do

The existing desk screen is a bibliographic and integrity preflight. It frequently detects unidentifiable references and uncheckable high-risk prose, but it does not reason over study design, methods, statistics, controls or claim support.

Scientific issue extractionnot_implemented
Arbitrary claim extractionnot_implemented
Review / manuscript alignmentnot_implemented
Revision-resolution judgementnot_implemented

The evidence-led build order

01

Extract claims and issues

Exact quote, locator, typed proposition, uncertainty and high-recall candidate generation.

02

Check study validity

Design, statistical unit, replication, controls, multiplicity, sensitivity and claim/design mismatch.

03

Track resolution

Align review issues to claims, compare versions semantically and require evidence for resolution.

Stages 01–03 plus revision comparison and dual-review/adjudication UI are now implemented. The next gate is to score expert pilot and development labels, freeze thresholds on validation, and open the sealed set once.

The product wedge

Biosingularity should be the evidence-risk control layer between literature or analysis and a consequential decision: what is claimed, what exact record supports it, what changed, and what is unsafe to assert.

Primary

Pharma & biotech evidence teams

Translational R&D, scientific affairs, evidence strategy, research integrity and R&D data/AI governance.

Secondary

Publishers, CROs & funders

Pre-submission checks, due diligence, editorial triage, systematic review and portfolio evidence monitoring.

Differentiation

A gate, not another search box

Exact claim receipts, fail-closed decisions, version-aware reverification and transparent peer-review evaluation.

Inspect the release evidence

No manuscript or review text is published. The repository contains the methods, schemas, aggregate report, corpus/split ids and prediction hashes.

Full technical findings →
Open the no-human proxy validation dashboard →
Expert review protocol →
medRxiv TDM policy → · eLife terms →