Orion

  • Scientific literature can operate as humanity's largest labeled dataset.
  • Past research contains latent value that was constrained by the computational limits, data availability, instrumentation, and theoretical context of its time.
  • Modern AI, computation, simulation, and knowledge graphs can recover that latent value.
  • The objective is to maximize the return on the world's existing scientific investment by identifying work that could effectively and productively be revisited today.
Documents analyzed

1,028,640
scorecards on record, any vocabulary epoch

browse as cards →

Spans found

803,739
vocabulary matches located

review these →

Spans admitted

142,436
classified to an inertia type

review these →

Vocabulary

1,576
surrender phrases · 1,564 buckets

Quality measured on the 100-paper curated calibration run, not claimed for the estate: precision 0.23, recall 0.82, F1 0.36 against a gold set of 5 papers, 767 sentences, 34 labelled constraints; 78,856 sentences judged, 9,204 candidates (11.7%) — calibration run only. A candidate is not a confirmed constraint.

Classification coverage

Of 803,739 located spans, 205,053 have been decided and 142,436 admitted. Coverage varies across strata, so admitted-span counts are not comparable between fields without correction.

Provisional scores

The six factors are rule-derived from the admitted evidence. A score is rule-provisional unless an analyst assigned it. Analyst-assigned scorecards in the corpus: 4.

Composite

The composite is the product of the six factors, not a sum: a zero on any factor is a zero overall.

Verdict counting

Verdict rows count distinct documents ever given each verdict at any adjudicator version, so per-verdict figures may sum above the adjudicated total.

every span verified against the stored bytes before admission · offsets index the doc_cid text

The funnel

1,560,771
Documents known — estate manifest inventory, incl. the curated deterministic draw (100% of known)
1,560,384
Documents retrieved — files present, content-addressed (100% of known)
1,028,640
Documents indexed — full text extracted and content-addressed (66% of known)
1,028,640
Documents analyzed — scorecards on record, any vocabulary epoch (66% of known)
142,436
Surrender spans admitted — spans, not documents: vocabulary matches classified to an inertia type (9% of known)
51
Constraint located (curated corpus, 96 screened) — a stated limit found at a verifiable offset
46
Survived adjudication (curated corpus) — constraint is scientific and relievable
4
Revival candidates (curated corpus) — capability match confirmed; test specifiable
1
Counterfactuals run (curated corpus) — recomputed on the PAF floor with modern methods

estate manifest orion-manifest · curated corpus cdb7c0caadb5 · estate bar widths are proportional to documents known, recomputed live (a visual minimum keeps tiny stages legible; shares under 1% are labeled); curated stage widths scaled for legibility — the counts are the truth

Estate inventory

manifest orion-manifest · live counts, nothing baked

Documents known

1,560,771
estate manifest rows

Documents retrieved

1,560,384
files present and reachable

Documents indexed

1,028,640
full text content-addressed

Documents analyzed

1,028,640
scorecards on record, any vocabulary epoch

Vocabulary size

1,576
surrender phrases, epoch 1786619206

IndexContentsCount
orion-manifestEvery document known to the estate, one row per document. 1,560,771
orion-retrievedDocuments whose file is present and content-addressed. 1,560,384
orion-documentsDocuments with extracted text. 1,028,640
orion-scorecardsScorecards, one per analysed document, any vocabulary epoch. 1,028,640
orion-hitsConstraint spans located in document text. 803,739
  — decidedLocated spans the classifier has ruled on. 205,053
  — admittedDecided spans assigned to one of seven inertia types. 142,436
orion-assessmentsDeep assessments recorded against documents. 98,054
orion-verdictsVerdict rows; distinct documents per verdict, any adjudicator version. 91,694
  — adjudicatedDocuments holding a verdict at the current adjudicator. 54,137
vocabularyMatchable trigger phrases in the controlled vocabulary. 1,128

Corpus composition

estate: 1,028,640 documents analyzed · live counts · curated pies labeled

Years

publication decade, estate-wide · live · the muted slice is the remainder of everything known

live composition loads here

Corpus composition

estate: 1,028,640 documents analyzed · live counts · curated pies labeled

Disciplines

discipline group, estate-wide · live · the muted slice is the remainder of everything known

live composition loads here

Corpus composition

estate: 1,028,640 documents analyzed · live counts · curated pies labeled

Analysis lifecycle

every document known to the estate manifest, partitioned · live

live composition loads here

Bottleneck classes — estate

inertia type over classified surrender spans, estate-wide · live · spans still awaiting classification are counted beside the pie, never drawn as a slice competing with the classified types

live composition loads here

Screening verdicts — curated corpus

verdict over screened papers · curated corpus only (96 screened) — verdicts do not exist estate-wide

No stated constraint: 45Held for review: 42Rejected with reason: 5Revival candidate: 496
No stated constraint4547%
Held for review4244%
Rejected with reason55%
Revival candidate44%
Bottleneck classes — curated corpus

among located constraints · curated corpus only (51 constraints, analyst-adjudicated)

Data unavailable (cost): 25Data unavailable (cohort): 14CPU / parallelism: 1251
Data unavailable (cost)2549%
Data unavailable (cohort)1427%
CPU / parallelism1224%

Corpus analytics

analytics load here

Corpus analytics are not being served yet — this panel appears when the analytics endpoint lands. Nothing here is baked.

Head to head — revival rate by field

loads live from /api/orion/discipline-headtohead

measuring cohorts live…

Top ranked right now — the live head of the ranking

Open the analysis cards — every ranked paper as a sortable tile →

The current top of the one live ranking, computed over every scored document in the estate at view time. These are the top papers — ordered by the calibrated composite, re-ranked as extraction and scoring proceed (tiebreaked ordering lands when the ranking API ships it). The curated screening examples further down the page are calibration material, not the head of this list.

loads live from /api/orion/scorecards · composite desc · re-ranks at every view

the live top of the ranking loads here — if the ranking API is unreachable this section stays empty rather than promoting a stale #1; the full ranking is in “All documents” below

First counterfactual, computed

GSE1297 — hippocampal expression in incipient Alzheimer's disease (Blalock et al. 2004, PNAS; data public in GEO). The doctrine's question made concrete: what would this analysis conclude with today's methods? Same data, same test statistic, modern multiple-testing discipline.

control n=9 vs incipient n=7 · 22,283 probes · limma eBayes, BH · computed on the PAF floor

Era practice

696
genes at unadjusted p < 0.05 — the readout a 2004 pipeline reports.

Modern practice

0
genes at FDR < 0.05 — and 0 at FDR < 0.10.

What it means

At this design and sample size the data cannot support single-gene claims. Recorded as a labeled case, not a verdict on the original paper — its primary analysis correlated expression with cognitive score, and concordance against its published table is the recorded next step.

The library — every relevant paper, scored

Curated corpus: papers still in contention — candidates and held — each tile isolating the language that led to the scoring beside a spider graph over the six ranking factors. Papers judged not worth pursuing are not retained here; they are recorded below with their reasons and where to get them. The live, estate-wide ranking is the “All documents” section below.

Expected societal impactProbability of successReproducibilityRemaining unexplored valueCross-disciplinary leverageTime and cost to validate

candidate 0.09 2005 · Biochemistry, Genetics and Mol
Expected societal impact: 0.45Probability of success: 0.85Reproducibility: 0.90Remaining unexplored value: 0.60Cross-disciplinary leverage: 0.50Time and cost to validate: 0.90
Binding Site Graphs: A New Graph Theoretical Framework for Prediction of Transcription Fac
“Constructing a graph for each motif width, however, would require the ensemble sampling procedure to be repeated many times (once for each width of interest). Doing so is computationally infeasible with available technology; we le”

Authors named the experiment they dropped: a motif-width sweep judged infeasible with available technology and left for a future study. Embarrassingly parallel; sequence data is public; runs on the PAF distributed floor as-is.

analyst-provisional · cc8be9cd83a4
candidate 0.08 2004 · Biochemistry, Genetics and Mol
Expected societal impact: 0.40Probability of success: 0.85Reproducibility: 0.95Remaining unexplored value: 0.55Cross-disciplinary leverage: 0.45Time and cost to validate: 0.95
Spatial Dimensions of Population Viability
“However, the exact calculation of q (which is obtained as the eigenvalue of a 2n − 1 by 2n − 1 matrix) becomes computationally prohibitive as the number of patches n grows.”

Exact metapopulation persistence requires the eigenvalue of a 2^(n-1) x 2^(n-1) matrix, abandoned as n grows. Self-contained; feasibility is arithmetic, and both verdicts (n=15 minutes, n=20 dense still infeasible) are recorded.

analyst-provisional · 62ab7a60c838
candidate 0.06 2004 · Biochemistry, Genetics and Mol
Expected societal impact: 0.45Probability of success: 0.80Reproducibility: 0.80Remaining unexplored value: 0.50Cross-disciplinary leverage: 0.50Time and cost to validate: 0.85
Gapped alignment of protein sequence motifs through Monte Carlo optimization of a hidden M
“Since it is computationally prohibitive for the sampler to consider many transitions at one time, a key design issue is the selection of allowed transitions between points. Results and discussion The block-motif model We first def”

The Gibbs sampler restricted its move set because considering many transitions at once was prohibitive. The unrestricted sampler is a direct, parallel re-run.

analyst-provisional · f994657a7f12
candidate 0.05 2005 · Biochemistry, Genetics and Mol
Expected societal impact: 0.60Probability of success: 0.70Reproducibility: 0.80Remaining unexplored value: 0.35Cross-disciplinary leverage: 0.55Time and cost to validate: 0.70
A method for aligning RNA secondary structures and its application to RNA motif detection
“Unfortunately, it is computationally prohibitive for even a moderate number of RNAs.”

Simultaneous RNA folding and alignment was prohibitive at the time and forced a two-stage approximation. The joint computation is now tractable at database scale.

analyst-provisional · ee1058588493
held 0.03 1986 · Biochemistry, Genetics and Mol
Expected societal impact: 0.60Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.65Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
Magnetic Field Measurements Aboard the JOIDES Resolution and Implications for Shipboard Pa
“Moreover, though it may be prohibitively expensive to shield the entire paleomagnetic laboratory, the results of this study suggest two ways in which the contamination of sensitive paleomagnetic samples can be reduced.”
rule-provisional · d38f3d9a21a2
held 0.03 1977 · Biochemistry, Genetics and Mol
Expected societal impact: 0.60Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.65Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
Synthesis and biological incorporation of icons into macromolecules for NMR study. Progres
“e geometry and ligand-forming properties of the reactive site, the apparent inefficiency of presently defined methods for selec- tive incorporatioh of amino acids into antibodies and the high cost of enriched compounds makes this ”
rule-provisional · 72ca060e4f69
held 0.03 1989 · Neuroscience
Expected societal impact: 0.55Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.65Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
Classical conditioning of model systems: A behavioral review
“Unfortunately, an equivalence in responding among these groups can only result in a failure to reject the null hypothesis and could reflect insufficient statistical power or a large degree of variability in the data.”
rule-provisional · 69051c264a01
held 0.03 1976 · Neuroscience
Expected societal impact: 0.55Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.65Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
An inexpensive integrated circuit intracranial stimulator
“The diode bridge enables the use of an inexpensive dc meter movement; ac microammeters are prohibitively expensive for this application.”
rule-provisional · f20c9115fa05
held 0.03 1972 · Neuroscience
Expected societal impact: 0.55Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.65Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
The General Automation 18/30 as a system for the general analysis and acquisition of data
“ents for a department. Another factor often overlooked is the fact that with a centralized system, the input/output devices such as disk drives, card read/punch modules, etc., are all shared by the various users. Such devices are ”
rule-provisional · ed451d79b8ee
held 0.03 1970 · Neuroscience
Expected societal impact: 0.55Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.65Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
Genetic factors in developmental hydrocephalus.
“including the province of Quebec, the Maritime provinces, eastern parts of Ontario, eastern area of the North West Territories, northern parts of the United States" even the Carribean. and Twenty-one cases were not se en because i”
rule-provisional · 136630640e02
held 0.02 1995 · Medicine
Expected societal impact: 0.60Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.55Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
Humeral epicondylitis among gas- and waterworks employees
“Consequently, these studies have had insufficient statistical power to detect a moderate effect (eg, a twofold increase in risk) for the dichotomous exposure variables corninonly reported.”
rule-provisional · 2c9ed603b811
held 0.02 1997 · Immunology and Microbiology
Expected societal impact: 0.55Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.55Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
The development of in vitro techniques to facilitate the study of hemopoiesis in the rainb
“wn as c-kit ligand or steel factor (Heyworth & Spooncer, 1993). Other cytokines may have modulatory effects on the hemopoietic process, but these effects are not as significant. Since the commercially available, purified growth fa”
rule-provisional · aa949978da2e
held 0.02 1996 · Immunology and Microbiology
Expected societal impact: 0.55Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.55Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
Cryptosporidiosis in Wisconsin: A case-control study of post-outbreak transmission
“We found no evidence that drinking municipal water was associated with post-outbreak cryptosporidiosis. Our study was limited by the small sample size and by our inability to interview many potential postoutbreak case-patients bec”
rule-provisional · 5b2613924bfe
held 0.02 1992 · Immunology and Microbiology
Expected societal impact: 0.55Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.55Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
Some Factors Influencing the Prognosis of Some Common Bovine Neurological Diseases With Pa
“In agricultural veterinary practice the cost of treatment is critical to the farming client and a prolonged period of high dose penicillin G is prohibitively expensive apart from the valuable dairy cow.”
rule-provisional · 2e05f1e8b371
held 0.02 2005 · Medicine
Expected societal impact: 0.60Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.50Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
A simplified high-throughput method for pyrethroid knock-down resistance (kdr) detection i
“Single Nucleotide Polymorphism (SNP) detection is problematic with simple PCR approaches, requiring the use of highly toxic reagents [13] or prohibitively expensive equipment.”
rule-provisional · 0df9a146ba4e
held 0.02 2005 · Medicine
Expected societal impact: 0.60Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.50Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
Long-term follow-up of breast cancer survivors with post-mastectomy pain syndrome
“Women also expressed willingness to participate in the study; indeed, a questionnaire response rate of 82% at 9 years postoperatively is remarkable. Our study is limited by the small sample size, the absence of data on preoperativ”
rule-provisional · f8513cbfbbe1
held 0.02 2005 · Biochemistry, Genetics and Mol
Expected societal impact: 0.60Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.50Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
Glucocorticoid Receptor-Dependent Gene Regulatory Networks
“Thus, to perform SACO on groups of treated and untreated animals, as we have done here, is prohibitively expensive and timeconsuming. Although this work has expanded the set of genes known to be regulated by the GR, the list remai”
rule-provisional · 7842a582672b
held 0.02 2005 · Biochemistry, Genetics and Mol
Expected societal impact: 0.60Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.50Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
Method for three-dimensional visualization of neurodegeneration in cupric-silver stained s
“While failure to detect differences between the acute, trigger, and control groups could be explained by the insensitivity of the method, alternatively, this may be a consequence of insufficient statistical power for discriminatin”

Constraint language present; extraction garbled. Held.

rule-provisional · 3e987ee84a50
held 0.02 2005 · Biochemistry, Genetics and Mol
Expected societal impact: 0.60Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.50Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
Thermodynamic calculations in biological systems
“Adiabatic mapping simplifies the computationally intractable problem of computing the energy landscape by isolating the significant degrees of freedom and minimizing with respect to the remainder.”

Full energy-landscape computation was avoided via adiabatic mapping. Computable now, but requires structure data — a data dependency gates it, not compute.

rule-provisional · e528c09a3a88
held 0.02 2005 · Biochemistry, Genetics and Mol
Expected societal impact: 0.60Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.50Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
Toward a Better Understanding of Human Prion Diseases
“The researchers also discovered that seasonal screening of all donations provided little additional clinical benefit and was prohibitively expensive, and that screening throughout the year provided no additional benefit in any setti”
rule-provisional · 9aa1d8c4138e
held 0.02 2005 · Medicine
Expected societal impact: 0.60Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.50Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
Case mix and outcomes for admissions to UK adult, general critical care units with chronic
“t the risk for ventilator-associated pneumonia has been estimated to be 3% per day [24], it would not be surprising to find increased hospital mortality associated with intubation, and it is possible that the Seneff and Afessa stu”
rule-provisional · 249b4a42f7d4
held 0.02 2005 · Biochemistry, Genetics and Mol
Expected societal impact: 0.60Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.50Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
Genome-Wide Associations of Gene Expression Variation in Humans
“This pool of 63 genes is therefore enriched by 44 genes that appear to have significant signals within 1 Mb of the gene. Finally, we assigned significance based on a FDR of q ¼ 0.05. As mentioned above, it was computationally prohib”
rule-provisional · 8fdf15c4be1e
held 0.02 2004 · Medicine
Expected societal impact: 0.60Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.50Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
Advanced Life Support in Obstetrics in Ecuador: Teaching the Teachers
“The mannequins are central to the course and prohibitively expensive for many host institutions.”
rule-provisional · e98a6fcf155b
held 0.02 2004 · Biochemistry, Genetics and Mol
Expected societal impact: 0.60Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.50Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
Genetic and environmental sources of egg size variation in the butterfly Bicyclus anynana
“f plasticity are usually lower than trait heritabilities (Scheiner, 1993; Scheiner and Yampolski, 1998), they are relatively difficult to detect and we cannot rule out that the absence of genetic variation in the plastic response ”
rule-provisional · e93b993ecfa4
held 0.02 2004 · Biochemistry, Genetics and Mol
Expected societal impact: 0.60Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.50Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
Visual functioning following prenatal exposure to organic solvents
“by Khattak et al (1999) that there is an increased risk of major birth defects among women who reported health symptoms related to solvent exposure. This type of modeling was used because access to airborne levels of organic solve”
rule-provisional · 50874a30360a
held 0.02 2003 · Biochemistry, Genetics and Mol
Expected societal impact: 0.60Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.50Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
Domain fusion analysis by applying relational algebra to protein sequence and domain datab
“One reason for this limitation is that the computation time becomes "prohibitively expensive" [18]. As a minimum, the database must have a sequence table (denoted by S) and a domain layout table (denoted by D) with some key attrib”
rule-provisional · 5c293e777b80
held 0.02 2002 · Medicine
Expected societal impact: 0.60Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.50Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
The pattern of recurrence of adenocarcinoma of the oesophago-gastric junction
“There are some conflicting reports suggesting that lymphadenectomy confers no survival benefit (Hulscher et al, 2001) however such studies are either too small with insufficient statistical power, not stratified by tumour stage or”

Constraint language present but two-column extraction garbles the sentence; held until extraction quality is fixed.

rule-provisional · 0748aadd5b7c
held 0.02 2001 · Medicine
Expected societal impact: 0.60Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.50Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
Improvements to parallel plate flow chambers to reduce reagent and cellular requirements
“Thus, small molecules such as chemokines, peptides or non-peptide organic compounds that previously would be prohibitively expensive to study can now be studied.”
rule-provisional · 7bfe54ef5c45
held 0.02 2001 · Biochemistry, Genetics and Mol
Expected societal impact: 0.60Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.50Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
Effect of Ploidy Elevation, Copy Number and Parent-Of-Origin on Transgene Expression in Po
“The production of potato from botanical seed, or true potato seed (TPS), is now allowing farmers to produce potato in many areas where certified seed cannot be grown locally or is prohibitively expensive to import.”
rule-provisional · 3e8ba2beb706
held 0.02 2001 · Biochemistry, Genetics and Mol
Expected societal impact: 0.60Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.50Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
Antimicrobial Resistance
“However, these methods are often prohibitively expensive for widespread implementation, especially in LMICs.”
rule-provisional · 6a95e234c518
held 0.02 1987 · Agricultural and Biological Sc
Expected societal impact: 0.45Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.65Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
Étude de valorisation agricole des boues provenant des stations d'épuration des eaux de la
“Landfills are becoming prohibitively expensive as land costs spiral upwards and solid \\'aste disposaI regulations tighten to protect our groundwater supplies from toxic liquids leaching out of the landfills. An economically and e”
rule-provisional · 07a808a2fe58
held 0.02 1982 · Agricultural and Biological Sc
Expected societal impact: 0.45Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.65Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
Methanol production from eucalyptus wood chips. Attachment VI. Florida's eucalyptus energy
“Moreover, nontarget species in plantdtions on the best s i t e s e i t h e r require prohibitively expensive competition control measures o r create an unmanageable plantation environment.”
rule-provisional · 2040556fe5b4
held 0.02 1973 · Agricultural and Biological Sc
Expected societal impact: 0.45Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.65Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
A Request for Journals
“s extremely difficult. In this connection we are writing to ask if any of your readers our collection would be able to help us in building up ofback numbers of psychiatric journals. To purchase back numbers through commercial chan”
rule-provisional · 304af214d05d
held 0.02 2005 · Neuroscience
Expected societal impact: 0.55Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.50Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
Morality as natural history
“Alternatively, 'charging' is a behaviour that both lowand high-RHP individuals can display, but that might prove prohibitively expensive for low-RHP individuals, in the sense that the cost that they incur outweighs any benefits th”
rule-provisional · 49b237fbe32c
held 0.02 2005 · Neuroscience
Expected societal impact: 0.55Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.50Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
Abstracts of the 13th European Conference on Eye Movements 2005
“Unfortunately, over the years the sale prices of books have become prohibitively expensive and book chapters have increasingly been given a low rating in comparison to publications in peer reviewed journals.”
rule-provisional · ab11eaf71aa4
held 0.02 2005 · Neuroscience
Expected societal impact: 0.55Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.50Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
Neural systems for recognising emotion from facial expressions
“The results of the current study may be limited by the small sample size of our groups, given that the number of pre-symptomatic and genetically tested Huntington’s disease gene carriers is generally small.”
rule-provisional · fb3f33e38304
held 0.02 2005 · Neuroscience
Expected societal impact: 0.55Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.50Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
An investigation of the influence of gender and menstrual cycle phase on metacognitive jud
“design modification, and the fact that not all subjects provided valid data for each measure, it is difficult to interpret these results definitively. Inspection of the effect sizes observed in Table 7 indicates that the current s”
rule-provisional · bc9983990bca
held 0.02 2004 · Neuroscience
Expected societal impact: 0.55Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.50Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
Hippocampal neurogenesis and spatial memory: postnatal neurogenesis in yellow-pine chipmun
“A statistical difference between adults (females 58 + 1g, males 54 + 1g) and juveniles (females 55 + 1g, males 50 + 2g) of the same sex may exist, but our low sample sizes provided insufficient statistical power to detect it (Tuke”
rule-provisional · 46d8c2b0000a
held 0.02 2002 · Immunology and Microbiology
Expected societal impact: 0.55Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.50Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
Strategies for preventing the spread of fish and shellfish diseases
“g the hulls of ves sels. The International Maritime Organisation (IMO) hases tablished some voluntary guidelines on ballast water handling although some of the suggested solutions are impractical from an operational viewpoint, bei”
rule-provisional · f65fafd51411
held 0.02 2001 · Immunology and Microbiology
Expected societal impact: 0.55Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.50Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
The search for poxviral recombinase
“id represent a random of assortment of "Am,"B", and "C"va8ants. 4.5.1 IdentEcation of Other Enzymes Capable of PCR-Based Cloning Reaction, and Optimization of hotocol =cation of vaccinia virus DNA polymerase is both tirne consumin”
rule-provisional · 8308814fc622
held 0.02 2001 · Immunology and Microbiology
Expected societal impact: 0.55Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.50Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
Potential Infectious Etiologies of Atherosclerosis: A Multifactorial Perspective
“Three prospective therapeutic trials have been reported, but all had insufficient statistical power to resolve the question (reviewed in 23,48).”
rule-provisional · ce1fe6c222da
held 0.02 2001 · Immunology and Microbiology
Expected societal impact: 0.55Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.50Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
Feline immunodeficiency virus model of maternal-fetal HIV-1 transmission
“This regimen reduces rates of mother-to-child HIV-1 transmission in women who do not breast-feed from about 3-in-12 births to l-in-12, a reduction of two-thirds [22]. Unfortunately, this regimen is prohibitively expensive for most”
rule-provisional · 1d7f1c3a4f83
held 0.02 1996 · Agricultural and Biological Sc
Expected societal impact: 0.45Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.55Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
The National Envirothon Environmental Literacy Evaluation
“Where it is not feasible to undertake separate evaluations, this assessment, combined with others, will add to the overall picture of quality environmental education programs and provide additional support to new and continuing En”
rule-provisional · 711502478296
held 0.02 2005 · Agricultural and Biological Sc
Expected societal impact: 0.45Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.50Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
Risk factors associated with herd-level exposure of cattle in Nebraska, North Dakota, and
“seroconversion is a more accurate measure of seasonal disease risk to specific cattle, seroprevalence is more closely aligned with requirements for export of cattle to Canada. Additionally, a study of seroconversion in mature cows”
rule-provisional · d47e8a3c312d
held 0.02 2005 · Chemistry
Expected societal impact: 0.45Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.50Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
Atom centered potentials for the description and the design of chemical compounds within d
“coupled perturbed Hartree-Fock) are desirable as benchmarks, but unfortunately, for larger systems, theoretical treatments beyond MP2 rapidly become computationally intractable. Specifically, the dispersion interaction between ben”
rule-provisional · 0a0571c68c57
held 0.02 2003 · Agricultural and Biological Sc
Expected societal impact: 0.45Probability of success: 0.50Reproducibility: 0.60Remaining unexplored value: 0.50Cross-disciplinary leverage: 0.50Time and cost to validate: 0.50
Characterization and Application of Semichemical-Based Attract-And-Kill to Suppress Cabbag
“Registration of new insecticides is now prohibitively expensive in many crops.”
rule-provisional · 0378ab7e8430

All documents — the estate, ranked live

Open the analysis cards — every ranked paper as a sortable tile →

Every analyzed document in the estate, ranked at request time by the composite — the product of the six ranking factors: Expected societal impact · Probability of success · Reproducibility · Remaining unexplored value · Cross-disciplinary leverage · Time and cost to validate. Each card draws the six factors as a spider graph, carries the provenance label, the admitted surrender evidence with its inertia type, and the document’s era and extraction record. The order is derived from the live index on every request and re-ranks as documents are re-scored; new scorecards appear here without this page changing.

Expected societal impactProbability of successReproducibilityRemaining unexplored valueCross-disciplinary leverageTime and cost to validate

ranking loads live from /api/orion/scorecards · composite desc

estate scores are rule-provisional and conservatively scaled; calibration to the analyst scale is in progress — until it lands, composites here are not comparable to the analyst-provisional composites in the curated sections above

live scorecards load here

Investigated — the people behind the top papers

Every primary author of every top-ranked paper, named, placed and sourced. A scorecard ranks documents; it says nothing about the people who wrote them. This panel is the other half: for each of the top papers, who its primary authors actually are, where they are now, what they trained in, and who paid for the work — each with a date and, where one exists, a link to the primary source the claim was read from. Author positions were verified against obituaries, institutional directories, ORCID employment records and theses, and explicitly not against bibliographic aggregators, which mis-identified people repeatedly during this investigation; the failures are listed below rather than hidden. Where an author could not be placed, the panel says so in those words. It does not guess, and it does not leave the line blank.

loads live from /api/orion/author-dossiers

loading the author investigation…

Worked cases — three papers taken end to end

Each paper states a constraint in its own words. The constraint is quoted below with its location in the paper, the modern capability is applied, and the result is placed beside the paper's own number. Every arXiv identifier resolves, so a reader can retrieve each paper and check the quotation against the source.

three of three completed · outcomes only · arXiv identifiers per the arXiv identifier scheme (References, Method)

GALEX ultraviolet emission from M dwarfs

arXiv:1509.03645 · Jones, D. O., & West, A. A. (2016) · The Astrophysical Journal 817, 1 · doi:10.3847/0004-637X/817/1/1

“stars appear to emit less than the approximate NUV continuum, but this is likely due to our uncertainties. For these stars, we may have detected only the UV continuum, but stellar distance and temperature uncertainties make this impossible to determine with these data.” Section 4 (Results), p. 7 of the arXiv PDF
Capability applied
Gaia DR3 parallaxes. Median distance uncertainty 11.0% to 0.052%, a factor of about 200.
Side by side
Stars below the modelled photospheric continuum, PMSU sample n = 345: 254 with the paper's distances, 255 with Gaia DR3 distances. The named correction moves one star.
Binding term
The blackbody photosphere approximation. Median shift 1.554 dex, against 0.062 dex for distance — 24.9 times larger.
Model-free check
7 of 7 Hα-active early-M dwarfs sit below the blackbody floor, as do 159 of 160 early-M dwarfs. Against the corrected floor, 0 of 78.
Verdict
DISSOLVED — the paper's activity fractions stand; its stated reason does not.

Kepler triples via eclipse timing

arXiv:1510.08272 · Borkovits, T., Hajdu, T., Sztakovics, J., Rappaport, S., Levine, A., Bíró, I. B., & Klagyivik, P. (2016) · Monthly Notices of the Royal Astronomical Society 455, 4136–4165 · doi:10.1093/mnras/stv2530

“Due to insufficient length of the available data, however, we were unable to obtain reasonable LTTE solutions for most of these ETVs and, therefore, they are not included in our sample.” Section 6.5, p. 29 of the arXiv PDF
Capability applied
TESS. Observational baseline extended from 4.03 to 15.42 years, a factor of 3.83.
Side by side
Of the 58 affected systems, 56 cross to at least one full outer period and 44 reach two or more — the class the authors called robust.
Internal control
The paper sorted its 160 solutions into three tables on the exact quantity it blamed. The class covering more than two outer periods was later revised by more than a factor of two in 6% of cases; the one-to-two-period class in 20%; the flagged sub-one-period class in 37%.
Verdict
HELD — the constraint was load-bearing and correctly diagnosed.

Hipparcos variable-star classification

arXiv:0712.2898 · Willemsen, P. G., & Eyer, L. (2007)

“Since these approaches are computationally rather expensive, one should probably favour one of the simpler criteria for this stage of the testing procedure.” Section 3.1, ‘PCA Stopping criteria’, p. 11
Baseline first
Reproduced at 64.05% against the published 62.4%, per-class Spearman 0.942.
Side by side
The declined methods — bootstrap and parallel analysis — select 2 components and score 52.93%. The scree choice of 14 components scores 63.56%. An exhaustive sweep puts the optimum at 15 of 51.
Cost of the declined method
1.91 s, against 17.31 s for the grid search the same paper ran.
Better data, measured separately
+15.6 percentage points from Gaia DR3.
Verdict
HELD, then DISSOLVED — the judgement was right, the stated reason was not.

How candidates are ranked

A surrender phrase is where the authors state, in their own words, that something stopped them. That admission is the evidence: a paper with no surrender phrase has nothing to revive, and a paper with several has stated its own constraints explicitly. Everything below — the classification, the verdict, the six factors and the final order — is derived from those admissions and from nothing else.

The six factors, scored 0–1 and drawn on every paper below as a spider graph:

IMPExpected societal impact
PRBProbability of success
REPReproducibility
UNXRemaining unexplored value
XDLCross-disciplinary leverage
TCVTime and cost to validate
1 · Locate

The sweep matches every trigger phrase against every indexed document. No threshold and no cap — every match becomes a candidate span.

2 · Classify

A closed-set judge reads each located span and admits or rejects it: is this a first-person admission by these authors, or a citation, an aspiration, or rhetoric? Every admitted span is assigned one of seven inertia types — technological, data, economic, methodological, disciplinary, paradigm, credibility. The type says what kind of wall they hit.

3 · Adjudicate

Chained closed-set tests: a concession; era-bound rather than fundamental; then a promotion vote requiring two agreeing judgments — that the authors were prevented, and that what prevented them is a recoverable general capability rather than something irreplaceable. One verdict per document: Revival candidate, Held for review, Rejected with reason, or No stated constraint. Every rejection keeps its reason. The quantum exclusion applies here.

4 · Score

The six factors are rule-derived from the admitted surrender-phrase evidence — how many admissions, how many distinct inertia types, era and discipline — then calibrated onto the analyst scale. The composite is their product, not a sum: a zero anywhere is a zero overall, so a discredited result cannot ride a high benefit score into the top band. Scores are rule-provisional unless an analyst assigned them — analyst-assigned scorecards in the corpus: 4.

5 · Rank

Composite descending, tie-broken by admitted-span count, then inertia-type breadth, then older era first. That is the whole ordering.

Key performance indicators

Residual knowledge recovered

16,313
labeled counterfactual cases

New predictive performance

1
models outperforming the original

Validated discoveries

1
reproduced and confirmed findings

Downstream scientific impact

0
tracked over the long horizon

Curated calibration screen — analyst-scored examples

Not the top of the ranking — the analyst-scored screening examples that calibrate it. Papers whose authors stated a limit that present capability removes. The quoted sentence is the language that led to the scoring, located in the source text and verifiable against it — offsets come from the document, never from a model. The live top of the ranking is above; the full live ranking is in “All documents” below.

curated corpus · analyst-provisional — analyst-assigned here, rule-derived in the library; the calibration loop replaces both with measured values

0.09 Binding Site Graphs: A New Graph Theoretical Framework for Prediction of Transcription Factor Binding Sites 2005 · Biochemistry, Genetics and Molecular Biology · cc8be9cd83a4
Expected societal impact: 0.45Probability of success: 0.85Reproducibility: 0.90Remaining unexplored value: 0.60Cross-disciplinary leverage: 0.50Time and cost to validate: 0.90
six ranking factors · composite is their product
Expected societal impact0.45
Probability of success0.85
Reproducibility0.90
Remaining unexplored value0.60
Cross-disciplinary leverage0.50
Time and cost to validate0.90
“Constructing a graph for each motif width, however, would require the ensemble sampling procedure to be repeated many times (once for each width of interest). Doing so is computationally infeasible with available technology; we leave a comprehensive analysis of this strategy for a future study. Ulti”

Authors named the experiment they dropped: a motif-width sweep judged infeasible with available technology and left for a future study. Embarrassingly parallel; sequence data is public; runs on the PAF distributed floor as-is.

Authors: Timothy E. Reddy, Charles DeLisi, Boris E. Shakhnovich (Boston University)

0.08 Spatial Dimensions of Population Viability 2004 · Biochemistry, Genetics and Molecular Biology · 62ab7a60c838
Expected societal impact: 0.40Probability of success: 0.85Reproducibility: 0.95Remaining unexplored value: 0.55Cross-disciplinary leverage: 0.45Time and cost to validate: 0.95
six ranking factors · composite is their product
Expected societal impact0.40
Probability of success0.85
Reproducibility0.95
Remaining unexplored value0.55
Cross-disciplinary leverage0.45
Time and cost to validate0.95
“However, the exact calculation of q (which is obtained as the eigenvalue of a 2n − 1 by 2n − 1 matrix) becomes computationally prohibitive as the number of patches n grows.”

Exact metapopulation persistence requires the eigenvalue of a 2^(n-1) x 2^(n-1) matrix, abandoned as n grows. Self-contained; feasibility is arithmetic, and both verdicts (n=15 minutes, n=20 dense still infeasible) are recorded.

Authors: Mats Gyllenberg, Ilkka Hanski, J.A.J. Metz (University of Turku)

0.06 Gapped alignment of protein sequence motifs through Monte Carlo optimization of a hidden Markov model 2004 · Biochemistry, Genetics and Molecular Biology · f994657a7f12
Expected societal impact: 0.45Probability of success: 0.80Reproducibility: 0.80Remaining unexplored value: 0.50Cross-disciplinary leverage: 0.50Time and cost to validate: 0.85
six ranking factors · composite is their product
Expected societal impact0.45
Probability of success0.80
Reproducibility0.80
Remaining unexplored value0.50
Cross-disciplinary leverage0.50
Time and cost to validate0.85
“Since it is computationally prohibitive for the sampler to consider many transitions at one time, a key design issue is the selection of allowed transitions between points. Results and discussion The block-motif model We first define the alignment model in precise mathematical”

The Gibbs sampler restricted its move set because considering many transitions at once was prohibitive. The unrestricted sampler is a direct, parallel re-run.

Authors: Andrew F. Neuwald, Jun S. Liu (Cold Spring Harbor Laboratory)

0.05 A method for aligning RNA secondary structures and its application to RNA motif detection 2005 · Biochemistry, Genetics and Molecular Biology · ee1058588493
Expected societal impact: 0.60Probability of success: 0.70Reproducibility: 0.80Remaining unexplored value: 0.35Cross-disciplinary leverage: 0.55Time and cost to validate: 0.70
six ranking factors · composite is their product
Expected societal impact0.60
Probability of success0.70
Reproducibility0.80
Remaining unexplored value0.35
Cross-disciplinary leverage0.55
Time and cost to validate0.70
“Unfortunately, it is computationally prohibitive for even a moderate number of RNAs.”

Simultaneous RNA folding and alignment was prohibitive at the time and forced a two-stage approximation. The joint computation is now tractable at database scale.

Authors: Jianghui Liu, Jason T.L. Wang, Jun Hu, Bin Tian (Rutgers, The State University of New Jersey)

Constraint type beats field

Constraint type predicts revival better than field does. Papers whose stated limit was hardware or data are adjudicated Revival candidate two to three times as often as papers whose limit was methodological. Measured in five fields; 95% intervals do not overlap in any of them. The lowest-rate field filtered by constraint type still exceeds the highest-rate field unfiltered.

measured at epoch 1787428054 · label source orion-hits.classified · Wilson intervals · denominator is distinct documents · remainder convention: ≥1 methodological and zero technological/data

fieldtechnological / datamethodological remainderseparation
rate95% CInrate95% CIn
genetics0.7692[0.6166, 0.8735]390.2577[0.1811, 0.3528]972.98×
bio_other0.7564[0.6506, 0.8381]780.3400[0.2841, 0.4007]2502.22×
bioinformatics0.7000[0.6041, 0.7811]1000.3187[0.2642, 0.3787]2512.20×
astro0.6469[0.6143, 0.6781]8580.3259[0.3062, 0.3462]2,1081.98×
logistics0.5706[0.4970, 0.6413]1770.2059[0.1809, 0.2335]9082.77×

Seven inertia types

Each admitted sentence carries one inertia type. Counts are live. Each row links to the review tool filtered to that type.

counts from the adjudicated label, orion-hits.classified · injected at request time

typesentenceswhat it means
Hardware, instruments, compute
technological
10,015The machine, sensor or processor they had could not do it. The commonest kind that time actually cures.
“because of the lacking renewal property, it is not possible to compute this distribution directly”
judge these →
The data did not exist
data
6,625The measurements had not been taken, the sample was too small, the archive did not exist yet. Cured when somebody collects it.
“Due to limited sample sizes, these samples have less-significant detections”
judge these →
Money
economic
917It could have been done; nobody could afford to. Cured by price falling, not by cleverness.
“undue emphasis on improvements in air quality may have left insufficient local funding”
judge these →
No method existed — or they chose not to
methodological
124,597The largest class by far, and the least useful. It mixes a genuine absence of technique with “beyond the scope of this paper”, which is a choice, not an obstacle.
“the ESPRI consortium cannot carry out their planned program in a timely way”
judge these →
The field could not absorb it
disciplinary
122The work crossed a boundary nobody was standing on.
“Giovanni Lipi is something of a guess, as I have not been able to identify this person”
judge these →
The prevailing framework forbade it
paradigm
54The idea did not fit what the field believed at the time.
“We simply cannot get around the classical schemas in our thinking.”
judge these →
Nobody believed them
credibility
106The result was sound and was disregarded.
“for the last week I have been unable to leave my room”
judge these →

References

418 references, each verified against the registrant’s record. Start with a question rather than a filter: the row of questions below goes straight to a set. Below that, the two facets worth browsing — the role a work plays for this project, and the subject lineage it was collected in — are open; date, access and how each entry was verified fold into one drawer. 129 entries are tagged as arguing against this project and 29 as contested, and both are one click away. A third open rail marks the meta-science — the 162 entries whose own subject is research about research, or about how research is measured; source type joins the provenance drawer.

418 entries in 20 lineage groups · 364 carry a resolvable DOI · titles print exactly as deposited, capitalisation included; where a registrant deposits only a start page, only a start page is printed · nothing here is cited from memory, and an entry that could not be verified was left out rather than guessed at · the 243 entries added at epoch 1788105118 were each resolved through resolve.1788105118.py and then re-fetched and compared field by field by reverify.1788105118.py before publication · built by build.1788105118.py from references.1788105118.json; 74 distinct tags across 7 facets, every one traceable to the field it came from — source type is the registrant’s own deposited type, and each meta-science tag names the lineage group, the journal of record, or the verbatim title string it was read from

Start here

Each question sets one filter and clears the rest. Nothing here is scored or ranked beyond the tags you can see: 92 of the 418 entries carry foundational or nearest system, and that is all this page means by a core.

418 matched of 418
Narrow by
Role 12 — why this project cites it
Subject 20 — the lineage it was collected in
Meta-science 7 — 162 entries whose subject is research about research
Date, access and verification 35
Source type
Decade
Access
How it was verified
  1. Limitation and constraint extraction from scientific text — the nearest prior artThe problem Orion works on is not new. These are the people who did it first, and what they found.
  2. 001 Teufel, S., & Moens, M. (2002). Summarizing Scientific Articles: Experiments with Relevance and Rhetorical Status. Computational Linguistics, 28(4), 409–445. doi:10.1162/089120102762671936 — The origin of argumentative zoning: classifying every sentence of a paper by its rhetorical role, including the authors’ own statement of weakness. Orion’s locate-then-classify split is this idea narrowed to one zone.
  3. 002 Ioannidis, J. P. A. (2007). Limitations are not properly acknowledged in the scientific literature. Journal of Clinical Epidemiology, 60(4), 324–329. doi:10.1016/j.jclinepi.2006.09.011 — States the failure directly — limitations are systematically under-reported. The counterpart caution to any count of admitted constraints, including the ones on this page.
  4. 003 Teufel, S., Siddharthan, A., & Batchelor, C. (2009). Towards discipline-independent argumentative zoning: evidence from chemistry and computational linguistics. Proceedings of the 2009 Conference on Empirical Methods in Natural Language Processing Volume 3 - EMNLP '09, 3, 1493. doi:10.3115/1699648.1699696 — AZ-II, the scheme extended past computational linguistics into chemistry — the evidence that rhetorical-zone labels transfer across disciplines, which is what Orion assumes when it sweeps a mixed-discipline estate. Crossref deposits a start page only.
  5. 004 Liakata, M., Teufel, S., Siddharthan, A., & Batchelor, C. (2010). Corpora for the Conceptualisation and Zoning of Scientific Papers. Proceedings of the Language Resources and Evaluation Conference. doi:10.63317/2di6e2euu3as — The CoreSC annotation scheme and its corpora — the alternative to AZ, built from the components of a scientific investigation rather than its rhetoric.
  6. 005 Liakata, M., Saha, S., Dobnik, S., Batchelor, C., & Rebholz-Schuhmann, D. (2012). Automatic recognition of conceptualization zones in scientific articles and two life science applications. Bioinformatics, 28(7), 991–1000. doi:10.1093/bioinformatics/bts071 — CoreSC applied and measured on real full text; the reference point for what automatic zone recognition actually scores, against which any claim about Orion’s classifier has to be read.
  7. 006 ter Riet, G., Chesley, P., Gross, A. G., et al. (2013). All That Glitters Isn’t Gold: A Survey on Acknowledgment of Limitations in Biomedical Studies. PLoS ONE, 8(11), e73623. doi:10.1371/journal.pone.0073623 — Measures how often authors acknowledge limitations at all, and how thinly. It sets the ceiling on recall: a constraint that was never written down cannot be located.
  8. 007 Hu, Y., & Wan, X. (2015). Mining and Analyzing the Future Works in Scientific Articles. arXiv:1507.02140. arxiv.org/abs/1507.02140 — Mining the future-work sentences authors write at the end of a paper — the nearest published cousin of Orion’s target, a self-stated sentence about what has not been done yet.
  9. 008 Prabhakaran, V., Hamilton, W. L., McFarland, D., & Jurafsky, D. (2016). Predicting the Rise and Fall of Scientific Topics from Trends in their Rhetorical Framing. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1170–1180. doi:10.18653/v1/P16-1111 — Rhetorical framing tracked over time predicts the rise and fall of a research topic — the evidence that how authors talk about their own work carries recoverable signal, which is Orion’s whole premise.
  10. 009 Kilicoglu, H., Rosemblat, G., Malički, M., & ter Riet, G. (2018). Automatic recognition of self-acknowledged limitations in clinical research literature. Journal of the American Medical Informatics Association, 25(7), 855–861. doi:10.1093/jamia/ocy038 — The first system built specifically to recognise self-acknowledged limitations in clinical papers — the nearest published ancestor of the Locate and Classify stages.
  11. 010 Kilicoglu, H., Rosemblat, G., Hoang, L., et al. (2021). Toward assessing clinical trial publications for reporting transparency. Journal of Biomedical Informatics, 116, 103717. doi:10.1016/j.jbi.2021.103717 — The same group reading trial publications for what the authors did and did not disclose — the wider frame Orion’s constraint extraction sits inside.
  12. 011 Lan, M., Cheng, M., Hoang, L., ter Riet, G., & Kilicoglu, H. (2024). Automatic categorization of self-acknowledged limitations in randomized controlled trial publications. Journal of Biomedical Informatics, 152, 104628. doi:10.1016/j.jbi.2024.104628 — The current state of the art on limitation-sentence detection: not just finding the sentence but categorising the kind of limitation, which is exactly what Orion’s seven inertia types attempt.
  13. Discourse structure and sentence-role classification in scientific textBefore a limitation sentence can be classified it has to be found, and finding it is a sentence-role problem. This group holds the schemes, corpora and models that label what each sentence of a paper is doing — background, method, result, conclusion — and the discourse structures that hold those sentences together. Orion’s Locate stage is a narrow instance of this task, so its accuracy ceiling is set here.
  14. 012 MANN, W. C., & THOMPSON, S. A. (1988). Rhetorical Structure Theory: Toward a functional theory of text organization. Text - Interdisciplinary Journal for the Study of Discourse, 8(3). doi:10.1515/text.1.1988.8.3.243 — The original account of text as a tree of relations between spans — the theoretical ancestor of every rhetorical-zone scheme Orion’s classifier inherits.
  15. 013 Teufel, S., Carletta, J., & Moens, M. (1999). An annotation scheme for discourse-level argumentation in research articles. Proceedings of the ninth conference on European chapter of the Association for Computational Linguistics -, 110. doi:10.3115/977035.977051 — The annotation scheme for discourse-level argumentation in research articles, set out in full before any classifier was built — the definition of the zone labels that every later scheme, Orion’s included, is a variant of. Crossref deposits a start page only.
  16. 014 Teufel, S., Siddharthan, A., & Tidhar, D. (2006). Automatic classification of citation function. Proceedings of the 2006 Conference on Empirical Methods in Natural Language Processing, 103–110. aclanthology.org/W06-1613/ — The earlier and finer-grained citation scheme — prior art for the same problem of locating one sentence and labelling what its author was doing with it.
  17. 015 Hirohata, K., Okazaki, N., Ananiadou, S., & Ishizuka, M. (2008). Identifying Sections in Scientific Abstracts using Conditional Random Fields. Proceedings of the Third International Joint Conference on Natural Language Processing: Volume-I. aclanthology.org/I08-1050/ — The pre-neural CRF baseline for section labelling of abstracts — it already exploited sentence position heavily, and stands as a reminder of how much of this task position alone solves without reading the language at all. The venue string is taken verbatim from the ACL Anthology BibTeX record, brace protection included.
  18. 016 Prasad, R., Dinesh, N., Lee, A., et al. (2008). The Penn Discourse TreeBank 2.0.. Proceedings of the Language Resources and Evaluation Conference. doi:10.63317/55sr8v9kexr7 — The general antecedent for annotating discourse relations by their explicit connectives — the scheme against which every scientific discourse corpus is defined or contrasted.
  19. 017 Guo, Y., Korhonen, A., Liakata, M., Silins, I., Sun, L., & Stenius, U. (2010). Identifying the Information Structure of Scientific Abstracts: An Investigation of Three Different Schemes. Proceedings of the 2010 Workshop on Biomedical Natural Language Processing, 99–107. aclanthology.org/W10-1913/ — Three annotation schemes for the information structure of abstracts compared head to head on the same corpus — evidence that how much accuracy gets reported depends heavily on which scheme is being predicted, not only on the classifier.
  20. 018 Spooren, W., & Degand, L. (2010). Coding coherence relations: Reliability and validity. Corpus Linguistics and Linguistic Theory, 6(2). doi:10.1515/cllt.2010.009 — Cited against this project. Reliability of coherence-relation coding falls sharply as the scheme grows finer, and the agreement figures that get published are often computed in ways that overstate it — a direct challenge to fine-grained rhetorical labelling of the kind Orion depends on.
  21. 019 Dernoncourt, F., Lee, J. Y., & Szolovits, P. (2017). Neural Networks for Joint Sentence Classification in Medical Paper Abstracts. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers, 694–700. doi:10.18653/v1/e17-2110 — The joint sentence-classification network for medical abstracts — the result that labelling a sentence sequence together beats labelling each sentence on its own.
  22. 020 Dernoncourt, F., & Lee, J. Y. (2017). PubMed 200k RCT: a Dataset for Sequential Sentence Classification in Medical Abstracts. arXiv:1710.06071. arxiv.org/abs/1710.06071 — The large sequential sentence classification dataset that made this task trainable at scale — the data source that fixes what the sentence-role labels are and how much material is needed to learn them. Published at IJCNLP 2017; the ACL Anthology record is I17-2052.
  23. 021 Jin, D., & Szolovits, P. (2018). Hierarchical Neural Networks for Sequential Sentence Classification in Medical Scientific Abstracts. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, 3100–3109. doi:10.18653/v1/d18-1349 — Hierarchical sequential labelling over the sentences of an abstract — the architecture a section-aware sentence classifier of Orion’s kind descends from.
  24. 022 Jurgens, D., Kumar, S., Hoover, R., McFarland, D., & Jurafsky, D. (2018). Measuring the Evolution of a Scientific Field through Citation Frames. Transactions of the Association for Computational Linguistics, 6, 391–406. doi:10.1162/tacl_a_00028 — Citation function classified at corpus scale and then used to track how a field changed — the demonstration that a sentence-level rhetorical label, applied broadly enough, yields a metascientific signal and not merely a parse.
  25. 023 Yang, A., & Li, S. (2018). SciDTB: Discourse Dependency TreeBank for Scientific Abstracts. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 444–449. doi:10.18653/v1/p18-2071 — Discourse dependency structure annotated over scientific abstracts — the resource behind the claim that a paper’s argument has a parseable shape rather than only a sequence of roles.
  26. 024 Cohan, A., Beltagy, I., King, D., Dalvi, B., & Weld, D. (2019). Pretrained Language Models for Sequential Sentence Classification. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 3691–3697. doi:10.18653/v1/d19-1383 — Sequential sentence classification with a pretrained encoder run over the whole abstract, so a sentence is labelled in the company of its neighbours — the method behind classifying a candidate sentence in its section context rather than alone.
  27. 025 Cohan, A., Ammar, W., van Zuylen, M., & Cady, F. (2019). Structural Scaffolds for Citation Intent Classification in Scientific Publications. Proceedings of the 2019 Conference of the North, 3586–3596. doi:10.18653/v1/n19-1361 — Citation intent classified with auxiliary structural tasks, on a dataset an order of magnitude larger than the earlier citation-function corpora — the demonstration that a coarser label set and more data beat a fine-grained scheme on the same sentence-level problem.
  28. 026 Li, X., Burns, G., & Peng, N. (2021). Scientific Discourse Tagging for Evidence Extraction. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, 2550–2562. doi:10.18653/v1/2021.eacl-main.218 — Discourse tagging built expressly to feed evidence extraction downstream — the closest published statement of the pipeline Orion runs: label the discourse role of each sentence first, then use those labels to find the claim. First circulated as arXiv:1909.04758.
  29. 027 Brack, A., Entrup, E., Stamatakis, M., Buschermöhle, P., Hoppe, A., & Ewerth, R. (2024). Sequential sentence classification in research papers using cross-domain multi-task learning. International Journal on Digital Libraries, 25(2), 377–400. doi:10.1007/s00799-023-00392-z — Sentence-role classification trained and measured across several scientific domains at once — the direct test of whether such a model survives being moved to another field, and the record of how much accuracy is lost when it is.
  30. Scientific information extraction — entities, relations, and the datasets that define the taskWhat the field can actually pull out of a paper, and the annotated corpora that decide what counts as success. These are the benchmarks Orion’s Classify stage is measured against, and several of them are here precisely because their reported scores are lower than a downstream pipeline would like.
  31. 028 Kim, J. D., Ohta, T., Tateisi, Y., & Tsujii, J. (2003). GENIA corpus—a semantically annotated corpus for bio-textmining. Bioinformatics, 19(suppl_1), i180–i182. doi:10.1093/bioinformatics/btg1023 — The annotated substrate on which biomedical text mining was built — the reference point for what a carefully constructed scientific annotation resource costs and what it makes possible.
  32. 029 Kim, J. D., Ohta, T., Pyysalo, S., Kano, Y., & Tsujii, J. (2009). Overview of BioNLP'09 shared task on event extraction. Proceedings of the Workshop on BioNLP Shared Task - BioNLP '09, 1. doi:10.3115/1572340.1572342 — The shared task that moved biomedical extraction from entities to structured events — prior art for extracting a statement with arguments rather than a bare mention, which is what a stated constraint is. Crossref deposits a start page only.
  33. 030 Riedel, S., Yao, L., & McCallum, A. (2010). Modeling Relations and Their Mentions without Labeled Text. Lecture Notes in Computer Science, 148–163. doi:10.1007/978-3-642-15939-8_10 — Cited against this project. Measures how often the distant-supervision assumption is simply false on real text, and has to model the exception explicitly to recover — the negative result underneath any plan to bootstrap limitation labels from a weak signal.
  34. 031 Gupta, S., & Manning, C. (2011). Analyzing the Dynamics of Research by Extracting Key Aspects of Scientific Papers. Proceedings of 5th International Joint Conference on Natural Language Processing, 1–9. aclanthology.org/I11-1001/ — Focus, technique and domain pulled from abstracts with dependency patterns — early prior art for the aspect-labelled sentence, and a record of what the pattern-matching route reached before neural tagging.
  35. 032 Bada, M., Eckert, M., Evans, D., et al. (2012). Concept annotation in the CRAFT corpus. BMC Bioinformatics, 13(1), 161. doi:10.1186/1471-2105-13-161 — Full-text articles annotated with concepts from nine ontologies — the evidence that full text, not the abstract, is where the annotation cost and most of the ambiguity actually sit.
  36. 033 Krallinger, M., Leitner, F., Rabal, O., Vazquez, M., Oyarzabal, J., & Valencia, A. (2015). CHEMDNER: The drugs and chemical names extraction challenge. Journal of Cheminformatics, 7(S1), S1. doi:10.1186/1758-2946-7-s1-s1 — The community evaluation of chemical and drug name recognition — a crisply defined entity type with large training data still tops out short of perfect, which is the ceiling any fuzzier target such as a stated limitation must be read against.
  37. 034 Augenstein, I., Das, M., Riedel, S., Vikraman, L., & McCallum, A. (2017). SemEval 2017 Task 10: ScienceIE - Extracting Keyphrases and Relations from Scientific Publications. arXiv:1704.02853. arxiv.org/abs/1704.02853 — The shared task that fixed keyphrase and relation extraction over scientific text as a comparable benchmark — and reported top scores well short of what a downstream pipeline can lean on.
  38. 035 Reimers, N., & Gurevych, I. (2017). Reporting Score Distributions Makes a Difference: Performance Study of LSTM-networks for Sequence Tagging. Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, 338–348. doi:10.18653/v1/d17-1035 — Cited against this project. Single-score comparisons of neural sequence taggers are not reproducible — the distributions over random seeds overlap enough that many published rankings do not survive a reseed, which applies to every number Orion reports for its own classifier.
  39. 036 Luan, Y., Ostendorf, M., & Hajishirzi, H. (2017). Scientific Information Extraction with Semi-supervised Neural Tagging. arXiv:1708.06075. arxiv.org/abs/1708.06075 — Semi-supervised neural tagging for scientific keyphrases, showing how much unlabelled in-domain text has to be thrown at the problem before a sequence tagger works on this material — the method precedent for tagging spans in papers rather than matching patterns.
  40. 037 Ammar, W., Groeneveld, D., Bhagavatula, C., et al. (2018). Construction of the Literature Graph in Semantic Scholar. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 3 (Industry Papers), 84–91. doi:10.18653/v1/n18-3011 — How the Semantic Scholar literature graph is actually built, extraction stage by extraction stage — the industrial precedent for the corpus layer Orion sits on.
  41. 038 Nye, B., Li, J. J., Patel, R., et al. (2018). A Corpus with Multi-Level Annotations of Patients, Interventions and Outcomes to Support Language Processing for Medical Literature. arXiv:1806.04185. arxiv.org/abs/1806.04185 — Several thousand abstracts annotated at token level for participants, interventions and outcomes — the largest demonstration that fine-grained clinical annotation can be crowdsourced, and a plain record of how far agreement falls when it is.
  42. 039 Gábor, K., Buscaldi, D., Schumann, A. K., QasemiZadeh, B., Zargayouna, H., & Charnois, T. (2018). SemEval-2018 Task 7: Semantic Relation Extraction and Classification in Scientific Papers. Proceedings of The 12th International Workshop on Semantic Evaluation, 679–688. doi:10.18653/v1/s18-1111 — The follow-on shared task on semantic relations between scientific concepts — the second measurement point on the same benchmark line, a year later.
  43. 040 Kang, D., Ammar, W., Dalvi, B., et al. (2018). A Dataset of Peer Reviews (PeerRead): Collection, Insights and NLP Applications. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), 1647–1661. doi:10.18653/v1/n18-1149 — Peer reviews and their accept-or-reject decisions released as data — the closest public record of experts writing down what is wrong with a paper, which is the signal Orion tries to find inside the paper itself.
  44. 041 Luan, Y., He, L., Ostendorf, M., & Hajishirzi, H. (2018). Multi-Task Identification of Entities, Relations, and Coreference for Scientific Knowledge Graph Construction. arXiv:1808.09602. arxiv.org/abs/1808.09602 — SciERC: entities, relations and coreference annotated together over abstracts from six fields — the dataset that defines what scientific information extraction is asked to produce.
  45. 042 Wadden, D., Wennberg, U., Luan, Y., & Hajishirzi, H. (2019). Entity, Relation, and Event Extraction with Contextualized Span Representations. arXiv:1909.03546. arxiv.org/abs/1909.03546 — DyGIE++: span representations from a pretrained encoder shared across entity, relation and event extraction — the general architecture any candidate-sentence extractor for scientific text has to beat or borrow from.
  46. 043 Lehman, E., DeYoung, J., Barzilay, R., & Wallace, B. C. (2019). Inferring Which Medical Treatments Work from Reports of Clinical Trials. arXiv:1904.01606. arxiv.org/abs/1904.01606 — Given a trial report, decide the direction of an outcome and point at the span of text that says so — the span-plus-verdict output shape Orion adopts for adjudication.
  47. 044 Lee, J., Yoon, W., Kim, S., et al. (2019). BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36(4), 1234–1240. doi:10.1093/bioinformatics/btz682 — The first widely used biomedical BERT — the demonstration that pre-training on domain text, not merely fine-tuning on it, moves scientific extraction scores.
  48. 045 Lev, G., Shmueli-Scheuer, M., Herzig, J., Jerbi, A., & Konopnicki, D. (2019). TalkSumm: A Dataset and Scalable Annotation Method for Scientific Paper Summarization Based on Conference Talks. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2125–2131. doi:10.18653/v1/p19-1204 — Training data for scientific summarisation built by aligning conference talks to their papers — a scalable annotation trick of the sort Orion needs if it is to label limitation sentences without hand-annotating a corpus.
  49. 046 Neumann, M., King, D., Beltagy, I., & Ammar, W. (2019). ScispaCy: Fast and Robust Models for Biomedical Natural Language Processing. Proceedings of the 18th BioNLP Workshop and Shared Task, 319–327. doi:10.18653/v1/w19-5034 — Practical sentence splitting, tokenisation and entity linking tuned for scientific text — the toolchain Orion’s segmentation stage depends on, and the reason its sentence boundaries are not naive.
  50. 047 Yasunaga, M., Kasai, J., Zhang, R., et al. (2019). ScisummNet: A Large Annotated Corpus and Content-Impact Models for Scientific Paper Summarization with Citation Networks. Proceedings of the AAAI Conference on Artificial Intelligence, 33(01), 7386–7393. doi:10.1609/aaai.v33i01.33017386 — A hand-annotated corpus of scientific papers summarised with their citing sentences — the precedent for treating what other papers say about a work as evidence about that work.
  51. 048 Luan, Y., Wadden, D., He, L., Shah, A., Ostendorf, M., & Hajishirzi, H. (2019). A General Framework for Information Extraction using Dynamic Span Graphs. arXiv:1904.03296. arxiv.org/abs/1904.03296 — The dynamic span graph that propagates coreference and relation information between spans — prior art for treating the document, not the sentence, as the unit of extraction.
  52. 049 Yao, Y., Ye, D., Li, P., et al. (2019). DocRED: A Large-Scale Document-Level Relation Extraction Dataset. arXiv:1906.06127. arxiv.org/abs/1906.06127 — The benchmark for relations that can only be recovered by reading across sentences — it sets the expectation for the cross-sentence reasoning needed when a constraint and its subject are stated apart, and the gap it reports between the best models and human annotators remains wide.
  53. 050 Hou, Y., Jochim, C., Gleize, M., Bonin, F., & Ganguly, D. (2019). Identification of Tasks, Datasets, Evaluation Metrics, and Numeric Scores for Scientific Leaderboards Construction. arXiv:1906.09317. arxiv.org/abs/1906.09317 — Task, dataset, metric and score pulled out of papers to build leaderboards automatically — the closest published analogue to the claim that one specific factual assertion can be extracted from a paper at scale.
  54. 051 Alt, C., Gabryszak, A., & Hennig, L. (2020). TACRED Revisited: A Thorough Evaluation of the TACRED Relation Extraction Task. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 1558–1569. doi:10.18653/v1/2020.acl-main.142 — Cited against this project. Re-annotating a standard relation extraction benchmark showed that a large share of what looked like model error was annotation error — evidence that ceiling scores on extraction corpora measure the labels as much as the language.
  55. 052 Bender, E. M., & Koller, A. (2020). Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 5185–5198. doi:10.18653/v1/2020.acl-main.463 — Cited against this project. A model trained on form alone does not thereby acquire meaning — the standing objection to reading a classifier’s sentence labels as though the system had understood the constraint the author stated.
  56. 053 Ethayarajh, K., & Jurafsky, D. (2020). Utility is in the Eye of the User: A Critique of NLP Leaderboards. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 4846–4853. doi:10.18653/v1/2020.emnlp-main.393 — Cited against this project. Leaderboard rankings discard cost, calibration and the preferences of whoever has to use the output — the argument that a ranked candidate list tuned to one metric can still be worthless to the analyst reading it.
  57. 054 Linzen, T. (2020). How Can We Accelerate Progress Towards Human-like Linguistic Generalization?. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 5210–5217. doi:10.18653/v1/2020.acl-main.465 — Cited against this project. Gains on an in-distribution test set do not measure the generalisation a deployed system needs — the caution against reading held-out accuracy on limitation sentences as capability over an unseen corpus.
  58. 055 Kardas, M., Czapla, P., Stenetorp, P., et al. (2020). AxCell: Automatic Extraction of Results from Machine Learning Papers. arXiv:2004.14356. arxiv.org/abs/2004.14356 — Results extraction from the tables and prose of machine-learning papers — a worked instance of the same staged pipeline shape, and of how much of the effort lands in layout and table parsing rather than in the classifier.
  59. 056 Saier, T., & Färber, M. (2020). unarXive: a large scholarly data set with publications’ full-text, annotated in-text citations, and links to metadata. Scientometrics, 125(3), 3085–3108. doi:10.1007/s11192-020-03382-z — arXiv full text with in-text citations resolved to their targets — one of the few corpora where Orion can read a claim and then follow the citation that is supposed to support it.
  60. 057 Jain, S., Zuylen, M. V., Hajishirzi, H., & Beltagy, I. (2020). SciREX: A Challenge Dataset for Document-Level Information Extraction. arXiv:2005.00512. arxiv.org/abs/2005.00512 — Document-level extraction over full papers, where end-to-end performance on salient entities and document-level relations collapses next to sentence-level extraction — direct evidence that document-scale extraction from scientific text is nowhere near reliable enough to be trusted unreviewed.
  61. 058 Taillé, B., Guigue, V., Scoutheeten, G., & Gallinari, P. (2020). Let’s Stop Incorrect Comparisons in End-to-end Relation Extraction!. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 3689–3701. doi:10.18653/v1/2020.emnlp-main.301 — Cited against this project. End-to-end relation extraction results are routinely compared under incompatible evaluation settings, so part of the reported state of the art is an artefact — a warning that the extraction numbers Orion builds on are not directly comparable to each other.
  62. 059 Wadden, D., Lin, S., Lo, K., et al. (2020). Fact or Fiction: Verifying Scientific Claims. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 7534–7550. doi:10.18653/v1/2020.emnlp-main.609 — Verifying a scientific claim against abstracts and returning the sentences that support or refute it — the evidence-selection pattern Orion reuses when it puts sentences in front of an analyst.
  63. 060 Wang, L. L., Lo, K., Chandrasekhar, Y., et al. (2020). CORD-19: The COVID-19 Open Research Dataset. arXiv. doi:10.48550/arXiv.2004.10706 — The pandemic corpus assembled in weeks and mined at scale — the worked case for what a full-text scientific corpus release makes possible, and how quickly. arXiv:2004.10706; the DOI is registered with DataCite.
  64. 061 Bowman, S. R., & Dahl, G. (2021). What Will it Take to Fix Benchmarking in Natural Language Understanding?. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 4843–4855. doi:10.18653/v1/2021.naacl-main.385 — Cited against this project. Sets out how comprehensively benchmark construction fails to track real progress, from annotation artefacts to unreliable test sets — the reason a strong score on a limitation-sentence benchmark is weak evidence that the pipeline works.
  65. 062 Elangovan, A., He, J., & Verspoor, K. (2021). Memorization vs. Generalization : Quantifying Data Leakage in NLP Performance Evaluation. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, 1325–1335. doi:10.18653/v1/2021.eacl-main.113 — Cited against this project. Quantifies train-test overlap in common datasets and the score inflation it produces — the check Orion’s own evaluation has to pass before its accuracy figures mean anything, given how formulaic limitation sentences are.
  66. 063 Gu, Y., Tinn, R., Cheng, H., et al. (2021). Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing. ACM Transactions on Computing for Healthcare, 3(1), 1–23. doi:10.1145/3458754 — The controlled study finding that pre-training from scratch on domain text beats mixed-domain continual pre-training, with the BLURB benchmark built to measure it — the evidence behind choosing a domain encoder for Orion.
  67. 064 Heddes, J., Meerdink, P., Pieters, M., & Marx, M. (2021). The Automatic Detection of Dataset Names in Scientific Articles. Data, 6(8), 84. doi:10.3390/data6080084 — Detecting dataset names in articles: a narrow extraction target of exactly Orion’s kind, and a measure of how brittle the task becomes once the things being named are not drawn from a closed list. Crossref deposits a single page value, the article number.
  68. 065 Hou, Y., Jochim, C., Gleize, M., Bonin, F., & Ganguly, D. (2021). TDMSci: A Specialized Corpus for Scientific Literature Entity Tagging of Tasks Datasets and Metrics. arXiv:2101.10273. arxiv.org/abs/2101.10273 — A corpus annotating task, dataset and metric mentions where they occur in running prose rather than in tables — the annotation precedent for tagging a technical claim at the sentence in which it is stated.
  69. 066 Şahinuç, F., Tran, T. T., Grishina, Y., Hou, Y., Chen, B., & Gurevych, I. (2024). Efficient Performance Tracking: Leveraging Large Language Models for Automated Construction of Scientific Leaderboards. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 7963–7977. doi:10.18653/v1/2024.emnlp-main.453 — The same leaderboard-construction problem attempted again with large language models, five years on — the current state of automated claim extraction from papers, and the yardstick for Orion’s own extraction stage.
  70. Getting text off the page — PDF extraction, layout and tablesEvery claim on this site begins as a character stream recovered from a PDF that was never meant to be read by a machine. This group is the measured quality of that recovery: extraction benchmarks, layout and reading-order models, table and figure extraction, and the studies that quantify how extraction noise propagates into everything downstream. It is the least glamorous dependency in the pipeline and the one most likely to cap it.
  71. 067 Bird, S., Dale, R., Dorr, B., et al. (2008). The ACL Anthology Reference Corpus: A Reference Dataset for Bibliographic Research in Computational Linguistics. Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC'08). aclanthology.org/L08-1005/ — The corpus released with its extracted text and citation structure alongside the original PDFs — the template for publishing an estate in a form other people can check. The ACL Anthology record for this paper carries no DOI.
  72. 068 Councill, I., Giles, C. L., & Kan, M. Y. (2008). ParsCit: an Open-source CRF Reference String Parsing Package. Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC'08). aclanthology.org/L08-1291/ — ParsCit, the CRF reference-string parser most later work is measured against; the prior art for turning a bibliography block into records that can be linked. The ACL Anthology record for this paper carries no DOI.
  73. 069 Lopez, P. (2009). GROBID: Combining Automatic Bibliographic Data Recognition and Term Extraction for Scholarship Publications. Lecture Notes in Computer Science, 473–474. doi:10.1007/978-3-642-04346-8_62 — GROBID, the machine-learned PDF-to-TEI extractor used to turn a downloaded paper into headed sections and a parsed bibliography before any sentence is looked at.
  74. 070 Lopresti, D. (2009). Optical character recognition errors and their effects on natural language processing. International Journal on Document Analysis and Recognition (IJDAR), 12(3), 141–151. doi:10.1007/s10032-009-0094-8 — Traces character errors forward into tokenisation, tagging and parsing and shows the damage compounds rather than cancels — the mechanism by which a poor scan quietly corrupts a limitation sentence.
  75. 071 Constantin, A., Pettifer, S., & Voronkov, A. (2013). PDFX. Proceedings of the 2013 ACM symposium on Document engineering, 177–180. doi:10.1145/2494266.2494271 — PDFX, the fully automated PDF-to-XML converter that predates machine-learned layout models — the shape of the problem when it was still solved with rules, and the baseline a rules-based extractor is measured against. Crossref deposits the short title without its subtitle. The registrant deposits the main title only; the subtitle after the colon is not in the record, so it is not printed here.
  76. 072 Lipinski, M., Yao, K., Breitinger, C., Beel, J., & Gipp, B. (2013). Evaluation of header metadata extraction approaches and tools for scientific PDF documents. Proceedings of the 13th ACM/IEEE-CS joint conference on Digital libraries, 385–386. doi:10.1145/2467696.2467753 — Header metadata extraction measured across seven tools and found well short of what an automated ingest quietly assumes — the reason title and author on an ingested paper are checked against a registry rather than believed.
  77. 073 Klampfl, S., Jack, K., & Kern, R. (2014). A Comparison of Two Unsupervised Table Recognition Methods from Digital Scientific Articles. D-Lib Magazine, 20(11/12). doi:10.1045/november14-klampfl — Two unsupervised table recognisers compared on real scientific articles, with neither reaching a rate that would let a downstream reader trust the output unattended — a caution against treating extracted structure as ground truth.
  78. 074 Tkaczyk, D., Szostek, P., Fedoryszak, M., Dendek, P. J., & Bolikowski, Ł. (2015). CERMINE: automatic extraction of structured metadata from scientific literature. International Journal on Document Analysis and Recognition (IJDAR), 18(4), 317–335. doi:10.1007/s10032-015-0249-8 — CERMINE, the open alternative for pulling structured metadata and body text out of an article; the comparison point against which whatever extractor the pipeline settles on has to be justified.
  79. 075 Clark, C., & Divvala, S. (2016). PDFFigures 2.0. Proceedings of the 16th ACM/IEEE-CS on Joint Conference on Digital Libraries, 143–152. doi:10.1145/2910896.2910904 — Figures and tables carry the numbers a limitation is often stated about; this is the tool that separates them from the running text, so a caption is never read as a sentence of prose. Crossref deposits the short title without its subtitle. The registrant deposits the main title only; the subtitle after the colon is not in the record, so it is not printed here.
  80. 076 Bast, H., & Korzen, C. (2017). A Benchmark and Evaluation for Text Extraction from PDF. 2017 ACM/IEEE Joint Conference on Digital Libraries (JCDL), 1–10. doi:10.1109/jcdl.2017.7991564 — The controlled benchmark that measured what PDF-to-text tools actually get wrong — word order, hyphenation, running heads, ligatures — and found no tool close to clean; the direct evidence that Orion’s Extract stage is a source of error rather than a solved step.
  81. 077 Chiron, G., Doucet, A., Coustaty, M., Visani, M., & Moreux, J. P. (2017). Impact of OCR Errors on the Use of Digital Libraries: Towards a Better Access to Information. 2017 ACM/IEEE Joint Conference on Digital Libraries (JCDL), 1–4. doi:10.1109/jcdl.2017.7991582 — Estimates the OCR error load of a national digital library and what it costs the people searching it — the estate-scale version of the problem, measured on real holdings rather than a test set.
  82. 078 Shigarov, A., Altaev, A., Mikhailov, A., Paramonov, V., & Cherkashin, E. (2018). TabbyPDF: Web-Based System for PDF Table Extraction. Communications in Computer and Information Science, 257–269. doi:10.1007/978-3-319-99972-2_20 — TabbyPDF, table extraction driven by the text-positioning information already inside the file rather than by rendering it to an image; the cheap route for born-digital papers, which is most of the estate.
  83. 079 Tkaczyk, D., Collins, A., Sheridan, P., & Beel, J. (2018). Machine Learning vs. Rules and Out-of-the-Box vs. Retrained. Proceedings of the 18th ACM/IEEE on Joint Conference on Digital Libraries, 99–108. doi:10.1145/3197026.3197048 — Reference-string parsing compared across learned and rule-based parsers, retrained and out of the box, and still short of reliable — the evidence that citation extraction is an open problem, not a library call. Crossref deposits the title without its subtitle. The registrant deposits the main title only; the subtitle after the colon is not in the record, so it is not printed here.
  84. 080 Westergaard, D., Stærfeldt, H. H., Tønsberg, C., Jensen, L. J., & Brunak, S. (2018). A comprehensive and quantitative comparison of text-mining in 15 million full-text articles versus their corresponding abstracts. PLOS Computational Biology, 14(2), e1005962. doi:10.1371/journal.pcbi.1005962 — Text mining over fifteen million full texts set against the matching abstracts, showing how much recoverable signal exists only in the body of a paper — the reason for paying the cost of PDF extraction at all.
  85. 081 Gao, L., Huang, Y., Dejean, H., et al. (2019). ICDAR 2019 Competition on Table Detection and Recognition (cTDaR). 2019 International Conference on Document Analysis and Recognition (ICDAR), 1510–1515. doi:10.1109/icdar.2019.00243 — The shared evaluation that fixes what “good” table recognition means, and whose scores show archival documents remain far harder than modern ones.
  86. 082 Zhong, X., Tang, J., & Jimeno Yepes, A. (2019). PubLayNet: Largest Dataset Ever for Document Layout Analysis. 2019 International Conference on Document Analysis and Recognition (ICDAR), 1015–1022. doi:10.1109/icdar.2019.00166 — PubLayNet, the large automatically labelled layout corpus that made document-layout segmentation trainable at scale; the data underneath every modern extractor Orion could adopt.
  87. 083 Gyawali, B., Anastasiou, L., & Knoth, P. (2020). Deduplication of Scholarly Documents using Locality Sensitive Hashing and Word Embeddings. Proceedings of the Twelfth Language Resources and Evaluation Conference, 901–910. aclanthology.org/2020.lrec-1.113/ — Near-duplicate detection over scholarly full text by hashing and embeddings; the mechanism that stops one paper entering the estate several times through different aggregators and being counted twice. The ACL Anthology record for this paper carries no DOI.
  88. 084 Hamdi, A., Jean-Caurant, A., Sidère, N., Coustaty, M., & Doucet, A. (2020). Assessing and Minimizing the Impact of OCR Quality on Named Entity Recognition. Lecture Notes in Computer Science, 87–101. doi:10.1007/978-3-030-54956-5_7 — Entity recognition falls away sharply as character accuracy drops, which is the failure mode nearest to Orion’s own span classifier working over noisy scanned papers.
  89. 085 Li, M., Cui, L., Huang, S., Wei, F., Zhou, M., & Li, Z. (2020). TableBank: Table Benchmark for Image-based Table Detection and Recognition. Proceedings of the Twelfth Language Resources and Evaluation Conference, 1918–1925. aclanthology.org/2020.lrec-1.236/ — TableBank — table detection and structure recognition supervised weakly from Word and LaTeX sources; the benchmark for the part of a page Orion must exclude rather than read as prose. The ACL Anthology record for this paper carries no DOI.
  90. 086 Li, M., Xu, Y., Cui, L., et al. (2020). DocBank: A Benchmark Dataset for Document Layout Analysis. Proceedings of the 28th International Conference on Computational Linguistics, 949–960. doi:10.18653/v1/2020.coling-main.82 — DocBank, token-level layout labels over half a million pages; the resource that lets a segmenter tell a section heading from a body sentence before any classifier sees either.
  91. 087 van Strien, D., Beelen, K., Ardanuy, M., Hosseini, K., McGillivray, B., & Colavizza, G. (2020). Assessing the Impact of OCR Quality on Downstream NLP Tasks. Proceedings of the 12th International Conference on Agents and Artificial Intelligence, 484–496. doi:10.5220/0009169004840496 — Measures downstream task degradation directly against OCR quality across several tasks; the quantified form of the warning that Orion’s recall is capped by the state of its input text.
  92. 088 Xu, Y., Li, M., Cui, L., Huang, S., Wei, F., & Zhou, M. (2020). LayoutLM: Pre-training of Text and Layout for Document Image Understanding. Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 1192–1200. doi:10.1145/3394486.3403172 — LayoutLM — text and two-dimensional position pretrained together, the model class that reads a page rather than a stream of characters.
  93. 089 Wang, Z., Xu, Y., Cui, L., Shang, J., & Wei, F. (2021). LayoutReader: Pre-training of Text and Layout for Reading Order Detection. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 4735–4744. doi:10.18653/v1/2021.emnlp-main.389 — Reading order learned rather than assumed; the step that decides whether a two-column paper yields sentences or interleaved fragments.
  94. 090 Huang, Y., Lv, T., Cui, L., Lu, Y., & Wei, F. (2022). LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking. Proceedings of the 30th ACM International Conference on Multimedia, 4083–4091. doi:10.1145/3503161.3548112 — LayoutLMv3, the unified text-and-image masking version; the current default for document understanding and the obvious upgrade path for the Extract and Segment stages.
  95. 091 Pfitzmann, B., Auer, C., Dolfi, M., Nassar, A. S., & Staar, P. (2022). DocLayNet: A Large Human-Annotated Dataset for Document-Layout Segmentation. Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 3743–3751. doi:10.1145/3534678.3539043 — DocLayNet, layout segmentation labelled by people rather than by conversion heuristics — the harder test set, and the one that shows how much automatically labelled layout data flatters a model.
  96. 092 Smock, B., Pesala, R., & Abraham, R. (2022). PubTables-1M: Towards comprehensive table extraction from unstructured documents. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 4624–4632. doi:10.1109/cvpr52688.2022.00459 — PubTables-1M, table detection, structure recognition and functional analysis over nearly a million tables from scientific articles; the scale at which table extraction stops being anecdotal.
  97. 093 Blecher, L., Cucurull, G., Scialom, T., & Stojnic, R. (2023). Nougat: Neural Optical Understanding for Academic Documents. arXiv. doi:10.48550/arXiv.2308.13418 — An end-to-end visual transformer that reads an academic page straight into markup, ignoring the PDF text layer entirely — the nearest current alternative to Orion’s extract-then-segment route. Registered with DataCite as arXiv:2308.13418; the arXiv API was rate-limiting at the time of resolution, so the DataCite deposit was read instead.
  98. 094 Meuschke, N., Jagdale, A., Spinde, T., Mitrović, J., & Gipp, B. (2023). A Benchmark of PDF Information Extraction Tools Using a Multi-task and Multi-domain Evaluation Framework for Academic Documents. Lecture Notes in Computer Science, 383–405. doi:10.1007/978-3-031-28032-0_31 — A multi-task benchmark across ten extraction tools that found none best at everything and all of them weak somewhere; the evidence that the choice of extractor silently changes what Orion is able to find.
  99. Hedging, speculation and negation — the linguistic substrateA constraint statement is a hedged, often negated, first-person claim. These define the language phenomenon the trigger vocabulary is trying to catch.
  100. 095 Hyland, K. (1996). Writing Without Conviction? Hedging in Science Research Articles. Applied Linguistics, 17(4), 433–454. doi:10.1093/applin/17.4.433 — The linguistics of scientific hedging — why a constraint statement is written the way it is, and why a phrase list keyed to that language is a plausible starting point.
  101. 096 Chapman, W. W., Bridewell, W., Hanbury, P., Cooper, G. F., & Buchanan, B. G. (2001). A Simple Algorithm for Identifying Negated Findings and Diseases in Discharge Summaries. Journal of Biomedical Informatics, 34(5), 301–310. doi:10.1006/jbin.2001.1029 — NegEx — the canonical simple regular-expression algorithm over clinical text, and the honest comparison point for a rule-and-lexicon design like Orion’s vocabulary sweep.
  102. 097 Light, M., Qiu, X. Y., & Srinivasan, P. (2004). The Language of Bioscience: Facts, Speculations, and Statements In Between. HLT-NAACL 2004 Workshop: Linking Biological Literature, Ontologies and Databases, 17–24. aclanthology.org/W04-3103 — Distinguishes fact from speculation from the ground in between; the annotation problem that Orion’s judge stage has to solve when it separates a genuine admission from an aspiration.
  103. 098 Medlock, B., & Briscoe, T. (2007). Weakly Supervised Learning for Hedge Classification in Scientific Literature. Proceedings of the 45th Annual Meeting of the Association of Computational Linguistics, 992–999. aclanthology.org/P07-1125 — Hedge classification learned from weak supervision rather than a hand-built list — an early demonstration that the learned route beats the lexicon on this exact phenomenon.
  104. 099 Szarvas, G. (2008). Hedge Classification in Biomedical Texts with a Weakly Supervised Selection of Keywords. Proceedings of ACL-08: HLT, 281–289. aclanthology.org/P08-1033/ — Hedge classification with weakly supervised keyword selection, reported together with the finding that cue lists do not carry over between corpora — the same lexicon-transfer problem a fixed trigger vocabulary faces.
  105. 100 Vincze, V., Szarvas, G., Farkas, R., Móra, G., & Csirik, J. (2008). The BioScope corpus: biomedical texts annotated for uncertainty, negation and their scopes. BMC Bioinformatics, 9(S11), S9. doi:10.1186/1471-2105-9-S11-S9 — The BioScope corpus: uncertainty and negation annotated with their scopes. Scope is why a matched trigger phrase is not yet a constraint — the phrase may be negated or attributed to someone else.
  106. 101 Farkas, R., Vincze, V., Móra, G., Csirik, J., & Szarvas, G. (2010). The CoNLL-2010 Shared Task: Learning to Detect Hedges and their Scope in Natural Language Text. Proceedings of the Fourteenth Conference on Computational Natural Language Learning — Shared Task, 1–12. aclanthology.org/W10-3001 — The shared task that put hedge detection and scope resolution on a common benchmark; the reference for what accuracy is achievable on this task and by what means.
  107. Self-admitted technical debt — the published precedent against a fixed lexiconThe closest methodological analogue outside science: engineers writing down their own limitations in code comments. This line abandoned fixed keyword lists for learned classifiers, which is direct published evidence against Orion’s current design.
  108. 102 Potdar, A., & Shihab, E. (2014). An Exploratory Study on Self-Admitted Technical Debt. 2014 IEEE International Conference on Software Maintenance and Evolution, 91–100. doi:10.1109/ICSME.2014.31 — The study that established self-admitted technical debt as a measurable phenomenon, using 62 hand-derived comment patterns — the fixed-lexicon design, stated plainly.
  109. 103 Bavota, G., & Russo, B. (2016). A large-scale empirical study on self-admitted technical debt. Proceedings of the 13th International Conference on Mining Software Repositories, 315–326. doi:10.1145/2901739.2901742 — The large-scale replication that quantified how much self-admitted debt fixed patterns actually surface across many projects — the scale evidence behind the same warning.
  110. 104 Maldonado, E. D. S., Shihab, E., & Tsantalis, N. (2017). Using Natural Language Processing to Automatically Detect Self-Admitted Technical Debt. IEEE Transactions on Software Engineering, 43(11), 1044–1062. doi:10.1109/TSE.2017.2654244 — Cited against this project. The same group abandoned the fixed pattern list because it generalised badly, replacing it with a learned classifier that outperformed it substantially. Orion currently runs a fixed trigger vocabulary; this is the published precedent that says that choice will cap its recall, and it is listed here rather than left out.
  111. Literature-based discoveryThe founding claim that value already sitting in the published record can be recovered by reading across it.
  112. 105 Swanson, D. R. (1986). Fish Oil, Raynaud’s Syndrome, and Undiscovered Public Knowledge. Perspectives in Biology and Medicine, 30(1), 7–18. doi:10.1353/pbm.1986.0087 — The founding demonstration: a treatable connection sat unread across two literatures that never cited each other. Orion’s claim that unrecovered value is already in the record starts here.
  113. 106 Swanson, D. R. (1986). Undiscovered Public Knowledge. The Library Quarterly, 56(2), 103–118. doi:10.1086/601720 — Names the object — knowledge that is public, published, and nevertheless undiscovered because nobody has put the pieces together.
  114. 107 Swanson, D. R. (1988). Migraine and Magnesium: Eleven Neglected Connections. Perspectives in Biology and Medicine, 31(4), 526–557. doi:10.1353/pbm.1988.0009 — The second worked case, showing the method was not a one-off, and setting the template of eleven independently traceable connections.
  115. 108 Swanson, D. R., & Smalheiser, N. R. (1997). An interactive system for finding complementary literatures: a stimulus to scientific discovery. Artificial Intelligence, 91(2), 183–203. doi:10.1016/S0004-3702(97)00008-8 — The move from a hand-worked example to a system a person can drive — the same transition Orion is making from curated screen to running pipeline.
  116. 109 Smalheiser, N. R., & Swanson, D. R. (1998). Using ARROWSMITH: a computer-assisted approach to formulating and assessing scientific hypotheses. Computer Methods and Programs in Biomedicine, 57(3), 149–153. doi:10.1016/S0169-2607(98)00033-9 — ARROWSMITH in use: how a computer-assisted approach actually formulates and then assesses a hypothesis, which is the shape of Orion’s Adjudicate stage.
  117. 110 Henry, S., & McInnes, B. T. (2017). Literature Based Discovery: Models, methods, and trends. Journal of Biomedical Informatics, 74, 20–32. doi:10.1016/j.jbi.2017.08.011 — The standard modern survey of LBD models and methods; the map of what has been tried, against which Orion’s approach can be placed.
  118. 111 Sebastian, Y., Siew, E.-G., & Orimaye, S. O. (2017). Emerging approaches in literature-based discovery: techniques and performance review. The Knowledge Engineering Review, 32, e12. doi:10.1017/S0269888917000042 — A performance-focused review — what LBD systems actually score, and how rarely they are evaluated in a way that transfers.
  119. 112 Smalheiser, N. R. (2017). Rediscovering Don Swanson:The Past, Present and Future of Literature-based Discovery. Journal of Data and Information Science, 2(4), 43–64. doi:10.1515/jdis-2017-0019 — Swanson’s collaborator on what the field got right and wrong in thirty years; the retrospective that keeps this page honest about what LBD has and has not delivered.
  120. 113 Gopalakrishnan, V., Jha, K., Jin, W., & Zhang, A. (2019). A survey on literature based discovery approaches in biomedical domain. Journal of Biomedical Informatics, 93, 103141. doi:10.1016/j.jbi.2019.103141 — Survey of LBD in the biomedical domain, the subject area that supplies most of Orion’s curated calibration screen.
  121. 114 Thilakaratne, M., Falkner, K., & Atapattu, T. (2019). A systematic review on literature-based discovery workflow. PeerJ Computer Science, 5, e235. doi:10.7717/peerj-cs.235 — A systematic review of LBD as a workflow rather than an algorithm — the closest published description of the staged pipeline Orion runs.
  122. 115 Tshitoyan, V., Dagdelen, J., Weston, L., et al. (2019). Unsupervised word embeddings capture latent knowledge from materials science literature. Nature, 571(7763), 95–98. doi:10.1038/s41586-019-1335-8 — Latent knowledge recovered from published text by unsupervised embeddings, years before it was discovered experimentally — the strongest modern evidence that the record holds more than its authors extracted from it.
  123. Funded programmes for machine reading of science — what was tried before, and how it wentOrion is not the first attempt to read the literature at scale and act on what it finds. This group records the programmes and systems that went first — what they built, what they were evaluated against, and where the evaluations say they fell short. Their published shortfalls are the most direct evidence available about what a project of this shape can and cannot deliver.
  124. 116 King, R. D., Rowland, J., Oliver, S. G., et al. (2009). The Automation of Science. Science, 324(5923), 85–89. doi:10.1126/science.1165620 — The robot scientist that formed hypotheses, ran the experiments and interpreted the results without a human in the loop — the existence proof at the far end of the automation argument, on a domain small enough to close the loop.
  125. 117 Poon, H., Quirk, C., DeZiel, C., & Heckerman, D. (2014). Literome: PubMed-scale genomic knowledge base in the cloud. Bioinformatics, 30(19), 2840–2842. doi:10.1093/bioinformatics/btu383 — Machine reading run over the whole of PubMed to build a queryable knowledge base of genomic relations — the industrial-scale precedent for extracting structured claims from a corpus the size of the one Orion sweeps.
  126. 118 Cohen, P. R. (2015). DARPA’s Big Mechanism program. Physical Biology, 12(4), 045008. doi:10.1088/1478-3975/12/4/045008 — The funded programme that set out to do what Orion does, at scale and with a budget: read the research literature by machine, assemble the mechanism it describes, and propose what to test next — written by the manager who ran it.
  127. 119 Wei, C. H., Peng, Y., Leaman, R., et al. (2016). Assessing the state of the art in biomedical relation extraction: overview of the BioCreative V chemical-disease relation (CDR) task. Database, 2016. doi:10.1093/database/baw032 — Cited against this project. A community shared task measuring what biomedical relation extraction actually scores against expert curation — the published ceiling that any claim about Orion’s reading accuracy has to be read against.
  128. 120 Gyori, B. M., Bachman, J. A., Subramanian, K., Muhlich, J. L., Galescu, L., & Sorger, P. K. (2017). From word models to executable models of signaling networks using automated assembly. Molecular Systems Biology, 13(11), MSB177651. doi:10.15252/msb.20177651 — INDRA: the layer that turns sentences read out of papers into typed, executable statements, with provenance kept back to the sentence — the design Orion copies when it keeps every ranked candidate tied to the sentence it came from.
  129. 121 Rougier, N. P., Hinsen, K., Alexandre, F., et al. (2017). Sustainable computational science: the ReScience initiative. PeerJ Computer Science, 3, e142. doi:10.7717/peerj-cs.142 — A journal that publishes only replications of computational work, and the editorial machinery to make that reviewable — the venue-level answer to the question of who checks whether a stated computational result still stands.
  130. 122 Valenzuela-Escárcega, M. A., Babur, Ö., Hahn-Powell, G., et al. (2018). Large-scale automated machine reading discovers new cancer-driving mechanisms. Database, 2018. doi:10.1093/database/bay098 — REACH run over the whole cancer literature — the clearest published demonstration that a rule-and-machine-learning reader can sweep a corpus at Orion’s scale and surface mechanism statements a human missed.
  131. 123 Pyarelal, A., Valenzuela-Escarcega, M. A., Sharp, R., et al. (2020). AutoMATES: Automated Model Assembly from Text, Equations, and Software. arXiv:2001.07295. arxiv.org/abs/2001.07295 — The DARPA ASKE attempt to read text, equations and source code together and assemble a model from all three — prior art for the claim that a paper’s prose alone is not enough to recover what was actually done.
  132. 124 Alipourfard, N., Arendt, B., Benjamin, D. M., et al. (2021). Systematizing Confidence in Open Research and Evidence (SCORE). Center for Open Science. doi:10.31235/osf.io/46mnb — The design document for DARPA SCORE, the programme that tried to put a credibility score on individual published claims at scale — the closest funded precedent for ranking claims the way Orion ranks constraints.
  133. 125 Peterson, M., Korves, T., Garay, C., Kozierok, R., & Hirschman, L. (2022). Final Report on MITRE Evaluations for the DARPA Big Mechanism Program. arXiv:2211.03943. arxiv.org/abs/2211.03943 — Cited against this project. The independent evaluator’s final report on Big Mechanism, which had to invent a “reference set” of each paper’s major findings precisely because no gold standard exists for what a paper claims — the same measurement problem Orion has, stated by the people paid to measure it.
  134. 126 Bachman, J. A., Gyori, B. M., & Sorger, P. K. (2023). Automated assembly of molecular mechanisms at scale from text mining and curated databases. Molecular Systems Biology, 19(5), MSB202211325. doi:10.15252/msb.202211325 — The same assembly pipeline taken to the full literature and merged with curated databases; the scale report that says how much of a machine-read corpus survives deduplication and belief scoring.
  135. 127 Si, C., Yang, D., & Hashimoto, T. (2024). Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers. arXiv:2409.04109. arxiv.org/abs/2409.04109 — Cited against this project. The one large blinded comparison of machine-generated research ideas against expert ones, and it is genuinely contested: reviewers rated the machine ideas more novel but less feasible, which is exactly the failure mode of a ranked list of “liftable” constraints nobody can act on.
  136. 128 Lu, C., Lu, C., Lange, R. T., Foerster, J., Clune, J., & Ha, D. (2024). The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery. arXiv:2408.06292. arxiv.org/abs/2408.06292 — The current generation of end-to-end automated research systems, which generate ideas, run code and write the paper; useful to Orion as the boundary case, because it shows what the pipeline produces when nobody adjudicates the output.
  137. 129 Skarlinski, M. D., Cox, S., Laurent, J. M., et al. (2024). Language agents achieve superhuman synthesis of scientific knowledge. arXiv:2409.13740. arxiv.org/abs/2409.13740 — A retrieval-and-reading agent evaluated against human experts on literature questions, with the evidence trail kept — the current best case for the reading stage Orion depends on, and a claim strong enough to need checking.
  138. Metascience — research waste, reproducibility, and whether a result is worth anythingWhy an abandoned line of work is worth revisiting at all, and how much of the published record does not hold.
  139. 130 Ioannidis, J. P. A. (2005). Why Most Published Research Findings Are False. PLoS Medicine, 2(8), e124. doi:10.1371/journal.pmed.0020124 — The reason a revived result must be re-tested rather than trusted; the base rate against which any recovered finding has to be discounted.
  140. 131 Turner, E. H., Matthews, A. M., Linardatos, E., Tell, R. A., & Rosenthal, R. (2008). Selective Publication of Antidepressant Trials and Its Influence on Apparent Efficacy. New England Journal of Medicine, 358(3), 252–260. doi:10.1056/nejmsa065779 — Registered trials matched against what was published, showing efficacy inflated by which results reached print — the sharpest single measurement of publication bias, and the reason a corpus sweep sees a censored sample.
  141. 132 Chalmers, I., & Glasziou, P. (2009). Avoidable waste in the production and reporting of research evidence. The Lancet, 374(9683), 86–89. doi:10.1016/S0140-6736(09)60329-9 — Quantifies avoidable waste in the production and reporting of research — the size of the pool Orion is claiming to fish in.
  142. 133 Begley, C. G., & Ellis, L. M. (2012). Raise standards for preclinical cancer research. Nature, 483(7391), 531–533. doi:10.1038/483531a — The industry-side finding that most landmark preclinical results did not reproduce; the commercial cost of an unreliable record.
  143. 134 Piwowar, H. A., & Vision, T. J. (2013). Data reuse and the open data citation advantage. PeerJ, 1, e175. doi:10.7717/peerj.175 — Evidence that data put back into circulation gets reused and cited; the mechanism by which a revived result would pay off.
  144. 135 Chalmers, I., Bracken, M. B., Djulbegovic, B., et al. (2014). How to increase value and reduce waste when research priorities are set. The Lancet, 383(9912), 156–165. doi:10.1016/S0140-6736(13)62229-1 — Waste that begins at the point research priorities are set; the part of the problem a ranking system claims to address.
  145. 136 Gelman, A., & Loken, E. (2014). The Statistical Crisis in Science. American Scientist, 102(6), 460. doi:10.1511/2014.111.460 — The garden of forking paths, stated plainly: results can be selected without any deliberate p-hacking, simply by choices made after seeing the data — cited against this project, because the constraint an author writes down is chosen the same way.
  146. 137 Glasziou, P., Altman, D. G., Bossuyt, P., et al. (2014). Reducing waste from incomplete or unusable reports of biomedical research. The Lancet, 383(9913), 267–276. doi:10.1016/S0140-6736(13)62228-X — Waste from reports that are incomplete or unusable; directly why 253 selected papers on this page could not be retrieved and are recorded rather than dropped.
  147. 138 Ioannidis, J. P. A., Greenland, S., Hlatky, M. A., et al. (2014). Increasing value and reducing waste in research design, conduct, and analysis. The Lancet, 383(9912), 166–175. doi:10.1016/S0140-6736(13)62227-8 — Waste created in design, conduct and analysis — the stage where a stated constraint usually originates.
  148. 139 Macleod, M. R., Michie, S., Roberts, I., et al. (2014). Biomedical research: increasing value, reducing waste. The Lancet, 383(9912), 101–104. doi:10.1016/S0140-6736(13)62329-6 — The framing paper of the Lancet series — increasing value and reducing waste as a single measurable objective.
  149. 140 Freedman, L. P., Cockburn, I. M., & Simcoe, T. S. (2015). The Economics of Reproducibility in Preclinical Research. PLOS Biology, 13(6), e1002165. doi:10.1371/journal.pbio.1002165 — Puts a dollar figure on irreproducibility in preclinical research — the economic case for recovering rather than repeating.
  150. 141 Kaplan, R. M., & Irvin, V. L. (2015). Likelihood of Null Effects of Large NHLBI Clinical Trials Has Increased over Time. PLOS ONE, 10(8), e0132382. doi:10.1371/journal.pone.0132382 — Large trials stopped finding effects once preregistration was required — the cleanest natural experiment on how much of a published effect is an artefact of analytic freedom, and a caution on trusting the effect a limitation sentence is attached to.
  151. 142 Nosek, B. A., Alter, G., Banks, G. C., et al. (2015). Promoting an open research culture. Science, 348(6242), 1422–1425. doi:10.1126/science.aab2374 — The TOP guidelines — the transparency commitments that make a claim on this page checkable rather than merely asserted.
  152. 143 Nuijten, M. B., Hartgerink, C. H. J., van Assen, M. A. L. M., Epskamp, S., & Wicherts, J. M. (2015). The prevalence of statistical reporting errors in psychology (1985–2013). Behavior Research Methods, 48(4), 1205–1226. doi:10.3758/s13428-015-0664-2 — statcheck: recompute every reported test from the numbers printed in the paper, over three decades of psychology, and roughly half the papers contain an inconsistency — the precedent for a machine that reads what authors wrote and checks it rather than believing it. Crossref records the online-first year 2015; the issue is dated 2016.
  153. 144 Open Science Collaboration (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716. doi:10.1126/science.aac4716 — The measured replication rate in a whole field — the empirical anchor for treating a published result as provisional until reproduced.
  154. 145 Anderson, C. J., Bahník, Š., Barnett-Cowan, M., et al. (2016). Response to Comment on “Estimating the reproducibility of psychological science”. Science, 351(6277), 1037–1037. doi:10.1126/science.aad9163 — The original authors’ reply; printed here beside the comment because the disagreement is live and quoting only one side of it would be advocacy.
  155. 146 Baker, M. (2016). 1,500 scientists lift the lid on reproducibility. Nature, 533(7604), 452–454. doi:10.1038/533452a — What working scientists themselves say about reproducibility; the survey behind the assumption that unpublished failure is widespread.
  156. 147 Collberg, C., & Proebsting, T. A. (2016). Repeatability in computer systems research. Communications of the ACM, 59(3), 62–69. doi:10.1145/2812803 — Cited against this project. The systematic attempt to build the code behind hundreds of computer systems papers, and the rate at which it could not be done — the same finding as the R-file study, in the discipline that builds pipelines like Orion’s.
  157. 148 Gilbert, D. T., King, G., Pettigrew, S., & Wilson, T. D. (2016). Comment on “Estimating the reproducibility of psychological science”. Science, 351(6277), 1037–1037. doi:10.1126/science.aad7243 — Cited against this project. The published rebuttal arguing the headline reproducibility estimate is an artefact of low-powered and infidelitous replications rather than a crisis — the fairest form of dissent, aimed at the framing this page relies on.
  158. 149 Ioannidis, J. P. A. (2016). Why Most Clinical Research Is Not Useful. PLOS Medicine, 13(6), e1002049. doi:10.1371/journal.pmed.1002049 — Separates research that is true from research that is useful; the distinction the six ranking factors are trying to encode.
  159. 150 Wilkinson, M. D., Dumontier, M., Aalbersberg, I. J., et al. (2016). The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data, 3(1), 160018. doi:10.1038/sdata.2016.18 — FAIR — the conditions under which the material behind an old paper can be found and re-run at all.
  160. 151 Munafò, M. R., Nosek, B. A., Bishop, D. V. M., et al. (2017). A manifesto for reproducible science. Nature Human Behaviour, 1(1), 0021. doi:10.1038/s41562-016-0021 — The consolidated programme of reforms; the standard this page holds itself to when it publishes counts, intervals and its own failures.
  161. 152 Camerer, C. F., Dreber, A., Holzmeister, F., et al. (2018). Evaluating the replicability of social science experiments in Nature and Science between 2010 and 2015. Nature Human Behaviour, 2(9), 637–644. doi:10.1038/s41562-018-0399-z — Replication measured on experiments published in the two highest-profile venues — evidence that prominence does not predict reliability, which is why Orion ranks on stated constraint rather than on citation.
  162. 153 Fanelli, D. (2018). Is science really facing a reproducibility crisis, and do we need it to?. Proceedings of the National Academy of Sciences, 115(11), 2628–2631. doi:10.1073/pnas.1708272114 — Cited against this project. Argues the crisis narrative is not supported by the bibliometric evidence and that the framing itself distorts what needs fixing — a live disagreement about the premise from which Orion’s whole case for research waste is built.
  163. 154 Hardwicke, T. E., Mathur, M. B., MacDonald, K., et al. (2018). Data availability, reusability, and analytic reproducibility: evaluating the impact of a mandatory open data policy at the journal Cognition. Royal Society Open Science, 5(8), 180448. doi:10.1098/rsos.180448 — Cited against this project. A mandatory open-data policy raised availability sharply and still left most target values un-reproducible without author help — the measured gap between an artefact existing and a claim being checkable.
  164. 155 Piwowar, H., Priem, J., Larivière, V., et al. (2018). The state of OA: a large-scale analysis of the prevalence and impact of Open Access articles. PeerJ, 6, e4375. doi:10.7717/peerj.4375 — Measures how much of the literature is actually reachable; the ceiling on any pipeline that must fetch full text, including this one.
  165. 156 Redish, A. D., Kummerfeld, E., Morris, R. L., & Love, A. C. (2018). Reproducibility failures are essential to scientific inquiry. Proceedings of the National Academy of Sciences, 115(20), 5042–5046. doi:10.1073/pnas.1806370115 — Cited against this project. Makes the case that a failed replication is normal boundary-finding rather than waste — which would mean many of the constraints Orion flags as liftable are simply the field working correctly.
  166. 157 Silberzahn, R., Uhlmann, E. L., Martin, D. P., et al. (2018). Many Analysts, One Data Set: Making Transparent How Variations in Analytic Choices Affect Results. Advances in Methods and Practices in Psychological Science, 1(3), 337–356. doi:10.1177/2515245917747646 — Cited against this project. Twenty-nine teams, one dataset, one question, and answers that disagreed in direction as well as size — if the finding is that unstable, a constraint stated about it is not a fixed thing waiting to be lifted.
  167. 158 Stodden, V., Seiler, J., & Ma, Z. (2018). An empirical analysis of journal policy effectiveness for computational reproducibility. Proceedings of the National Academy of Sciences, 115(11), 2584–2589. doi:10.1073/pnas.1708290115 — Cited against this project. Even under a mandatory data and code policy, most requested artefacts could not be obtained and few results could be reproduced — evidence that the deposit behind a paper is usually not there to check the constraint against.
  168. 159 Botvinik-Nezer, R., Holzmeister, F., Camerer, C. F., et al. (2020). Variability in the analysis of a single neuroimaging dataset by many teams. Nature, 582(7810), 84–88. doi:10.1038/s41586-020-2314-9 — The same demonstration in neuroimaging, where no two independent pipelines produced the same map — cited against this project, because it puts a floor of analytic noise under any claim that a published limitation is now liftable.
  169. 160 Bird, A. (2021). Understanding the Replication Crisis as a Base Rate Fallacy. The British Journal for the Philosophy of Science, 72(4), 965–993. doi:10.1093/bjps/axy051 — Cited against this project. A philosophical account on which low replication rates follow from low prior odds in young fields and imply no misconduct or waste at all — the deepest available challenge to reading a stated limitation as a defect.
  170. 161 Errington, T. M., Mathur, M., Soderberg, C. K., et al. (2021). Investigating the replicability of preclinical cancer biology. eLife, 10, e71601. doi:10.7554/eLife.71601 — The preclinical cancer biology replication project, including how much published work could not even be attempted for lack of method detail — the same wall Orion hits when a paper cannot be retrieved or its method cannot be re-run.
  171. 162 Errington, T. M., Denis, A., Perfito, N., Iorns, E., & Nosek, B. A. (2021). Challenges for assessing replicability in preclinical cancer biology. eLife, 10, e67995. doi:10.7554/elife.67995 — Cited against this project. The companion paper reporting that not one of the sampled experiments could be designed from the published description alone, and every protocol needed the original authors — the sharpest evidence that a paper does not contain what a reader needs.
  172. 163 Schweinsberg, M., Feldman, M., Staub, N., et al. (2021). Same data, different conclusions: Radical dispersion in empirical results when independent analysts operationalize and test the same hypothesis. Organizational Behavior and Human Decision Processes, 165, 228–249. doi:10.1016/j.obhdp.2021.02.003 — Cited against this project. Dispersion this wide arose from how analysts operationalised the hypothesis, not from their statistics — which is exactly the step Orion cannot see when it reads a limitation sentence.
  173. 164 Trisovic, A., Lau, M. K., Pasquier, T., & Crosas, M. (2022). A large-scale study on research code quality and execution. Scientific Data, 9(1), 60. doi:10.1038/s41597-022-01143-6 — Cited against this project. Tens of thousands of deposited R files were run automatically and most of them failed — hard evidence that what a paper says was done cannot be recovered from its artefacts, let alone from its prose.
  174. 165 Jimenez, C. E., Yang, J., Wettig, A., et al. (2023). SWE-bench: Can Language Models Resolve Real-World GitHub Issues?. arXiv:2310.06770. arxiv.org/abs/2310.06770 — Grades by running the project’s own test suite instead of judging the diff, the cleanest demonstration that a pre-existing artefact can supply the ground truth — which is why Orion prefers a stated constraint in the authors’ own words to a label of its own manufacture.
  175. 166 Mialon, G., Fourrier, C., Swift, C., Wolf, T., LeCun, Y., & Scialom, T. (2023). GAIA: a benchmark for General AI Assistants. arXiv:2311.12983. arxiv.org/abs/2311.12983 — Established the short unambiguous final answer as the unit of agent evaluation, a convention this project deliberately does not adopt because Orion’s output is a ranked candidate that a human still has to adjudicate.
  176. 167 Liu, X., Yu, H., Zhang, H., et al. (2023). AgentBench: Evaluating LLMs as Agents. arXiv:2308.03688. arxiv.org/abs/2308.03688 — The first broad multi-environment agent suite, and the origin of the interaction trace as an evaluation unit — prior art for the question of what a staged pipeline should actually be scored on.
  177. 168 Gu, K., Shang, R., Jiang, R., et al. (2024). BLADE: Benchmarking Language Model Agents for Data-Driven Science. Findings of the Association for Computational Linguistics: EMNLP 2024, 13936–13971. doi:10.18653/v1/2024.findings-emnlp.815 — Scores the analytical decisions an agent makes rather than the answer it lands on, by comparing them against decisions expert analysts actually took — the design Orion follows when it asks a human to adjudicate the evidence instead of the verdict.
  178. 169 Laurent, J. M., Janizek, J. D., Ruzo, M., et al. (2024). LAB-Bench: Measuring Capabilities of Language Models for Biology Research. arXiv:2407.10362. arxiv.org/abs/2407.10362 — Separates the literature-reading tasks of biology research from the wet-lab ones and measures each, which is the split Orion depends on when it claims a text pipeline can say anything useful about a laboratory constraint.
  179. 170 Xie, T., Zhang, D., Chen, J., et al. (2024). OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments. Advances in Neural Information Processing Systems 37, 52040–52094. doi:10.52202/079017-1650 — Execution-based scoring in a real operating system rather than a simulated one, and the demonstration that scripted verification of an open-ended task is possible at all — the precondition for automating any check on Orion’s output.
  180. 171 Siegel, Z. S., Kapoor, S., Nadgir, N., Stroebl, B., & Narayanan, A. (2024). CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark. arXiv:2409.11363. arxiv.org/abs/2409.11363 — Turns computational reproducibility into a scored task an agent can be measured on, using published papers with their own code and data — the closest existing benchmark to what Orion would have to survive if its ranked candidates were ever checked automatically.
  181. 172 Chen, Z., Chen, S., Ning, Y., et al. (2024). ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery. arXiv:2410.05080. arxiv.org/abs/2410.05080 — Derives executable data-analysis tasks from the published literature and scores the program that results, which is the extraction-to-execution path Orion would need if a ranked candidate were ever to be tested rather than read.
  182. 173 Starace, G., Jaffe, O., Sherburn, D., et al. (2025). PaperBench: Evaluating AI’s Ability to Replicate AI Research. arXiv:2504.01848. arxiv.org/abs/2504.01848 — Grades replication of whole papers against author-written rubrics rather than a single number, and reports that the best agents fall well short of the researchers they are replicating — evidence against the assumption that a pipeline which can read a paper can also redo it.
  183. 174 Bragg, J., D'Arcy, M., Balepur, N., et al. (2025). AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite. arXiv:2510.21652. arxiv.org/abs/2510.21652 — A research-agent suite that insists on a held-constant harness and reported cost alongside the score, which is the reporting discipline the rest of this literature mostly lacks.
  184. 175 Phan, L., Gatti, A., Han, Z., et al. (2025). Humanity’s Last Exam. arXiv:2501.14249. arxiv.org/abs/2501.14249 — The reference point for expert-knowledge exams with checkable answers, cited here for what it does not measure — a system can answer hard closed questions and still be unable to carry a multi-stage task to a usable end.
  185. 176 Mitchener, L., Laurent, J. M., Andonian, A., et al. (2025). BixBench: a Comprehensive Benchmark for LLM-based Agents in Computational Biology. arXiv:2503.00096. arxiv.org/abs/2503.00096 — Open-ended bioinformatics analyses built from real published studies, on which frontier agents score close to the multiple-choice guessing rate — a hard limit on how much an automated reading of a paper can be trusted.
  186. 177 Sun, Q., Liu, Z., Ma, C., et al. (2025). ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows. arXiv:2505.19897. arxiv.org/abs/2505.19897 — Puts agents in front of the actual scientific software rather than a text description of it, which is the gap between reading that a constraint exists and being able to do anything about it.
  187. 178 Fa, D., Culjak, M., Pandza, B., & Cupic, M. (2026). BioAgent Bench: An AI Agent Evaluation Suite for Bioinformatics. arXiv:2601.21800. arxiv.org/abs/2601.21800 — An end-to-end bioinformatics suite whose unit of assessment is the produced artefact rather than a returned value, one of the few places the completeness of a scientific handoff is scored at all.
  188. 179 Su, L., Feng, Z., Chen, Z., et al. (2026). FrontierChallenge: Evaluating Scientific Workflow Completion. arXiv:2608.24979. arxiv.org/abs/2608.24979 — Three hundred end-to-end scientific workflows scored on whether the whole deliverable bundle arrives, not whether the answer looks right — the best configuration finished 20 of 97 tasks while averaging 87.9 out of 100, which is the clearest published statement that a high partial score does not mean the work was done. The source of the reference list ingested in this round.
  189. 180 Nair, S., Gunsalus, L., Orcutt-Jahns, B., et al. (2026). Agentic systems are adept at solving well-scoped, verifiable problems in computational biology. openRxiv. doi:10.64898/2026.04.06.716850 — Finds competence concentrated in problems that are already well scoped and mechanically verifiable, which is a warning about generalising from Orion’s tidy worked cases to the messy corpus it actually runs on.
  190. 181 Qu, Y., Lu, Y., Tu, X., et al. (2026). BiomniBench: Process-level Evaluation of LLM Agents for Real-world Biomedical Research. openRxiv. doi:10.64898/2026.05.12.724604 — Scores the process an agent followed rather than the endpoint it reached, on the argument that a right answer by a wrong route will not generalise — the same argument for showing an analyst the sentence and not just the rank.
  191. 182 Liu, T., Wang, A. X., Panescu, A., et al. (2026). Benchmarking AI Agents for Addressing Scientific Challenges Across Scales. arXiv:2606.12736. arxiv.org/abs/2606.12736 — Spans scientific problems from the molecular to the systemic and reports that difficulty does not track scale in the way practitioners assume, which cautions against any single corpus-wide difficulty prior.
  192. 183 Sun, Y., Han, X., Zhang, W., et al. (2026). Agents' Last Exam. arXiv:2606.05405. arxiv.org/abs/2606.05405 — Long-horizon professional work scored as delivered output, the nearest general-purpose analogue to asking whether a staged pipeline finished rather than whether it started well.
  193. 184 Shen, Y., Yang, Y., Xi, Z., et al. (2026). SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents. arXiv:2602.12984. arxiv.org/abs/2602.12984 — Multi-step scientific tool use scored step by step, useful here because Orion’s own failures are staged — extraction, segmentation, classification, calibration — and a single end-to-end number would hide which stage broke.
  194. 185 Wang, Y., Cheng, L., Zuo, Y., et al. (2026). NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers?. arXiv:2606.24530. arxiv.org/abs/2606.24530 — Asks whether an agent can reach the published state of the art of a Nature-family paper from the paper alone — the same question Orion asks about a stated limitation, posed as a measurement rather than an assumption.
  195. Claim verification, evidence retrieval and predicting what will replicateThe step past extraction: deciding whether a stated claim is supported, and whether a published finding would hold up if it were run again. This group carries both the benchmarks for evidence-backed verification and the literature on predicting replication from text — including the results that say such predictions do not generalise as well as their headline accuracies suggest.
  196. 186 Dreber, A., Pfeiffer, T., Almenberg, J., et al. (2015). Using prediction markets to estimate the reproducibility of scientific research. Proceedings of the National Academy of Sciences, 112(50), 15343–15347. doi:10.1073/pnas.1516179112 — Working researchers, priced in a market, already forecast which findings will fail — cited against this project, because it means much of what an automated ranker surfaces was known to the field before the machine read anything.
  197. 187 Thorne, J., Vlachos, A., Christodoulopoulos, C., & Mittal, A. (2018). FEVER: a Large-scale Dataset for Fact Extraction and VERification. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), 809–819. doi:10.18653/v1/n18-1074 — The dataset that defined claim verification as retrieve-evidence-then-decide, and the task formulation Orion’s adjudication step inherits when it asks an analyst to judge a candidate against the sentences behind it.
  198. 188 Altmejd, A., Dreber, A., Forsell, E., et al. (2019). Predicting the replicability of social science lab experiments. PLOS ONE, 14(12), e0225826. doi:10.1371/journal.pone.0225826 — Replication outcomes predicted from the reported statistics and study features rather than from the prose — the honest baseline any text-only ranker, Orion’s included, has to beat before its reading is doing the work.
  199. 189 Schuster, T., Shah, D., Yeo, Y. J. S., Roberto Filizzola Ortiz, D., Santus, E., & Barzilay, R. (2019). Towards Debiasing Fact Verification Models. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 3417–3423. doi:10.18653/v1/d19-1341 — Cited against this project. Shows that fact-verification models score well by reading giveaway words in the claim and largely ignoring the evidence — the standing warning that a sentence classifier trained on cue phrases will look accurate and be learning the wrong thing.
  200. 190 Gordon, M., Viganola, D., Bishop, M., et al. (2020). Are replication rates the same across academic fields? Community forecasts from the DARPA SCORE programme. Royal Society Open Science, 7(7), 200566. doi:10.1098/rsos.200566 — SCORE’s community forecasts across fields — the programme’s own evidence on how far expected replicability varies by discipline, which sets how much a single cross-discipline confidence scale can mean.
  201. 191 Pawel, S., & Held, L. (2020). Probabilistic forecasting of replication studies. PLOS ONE, 15(4), e0231416. doi:10.1371/journal.pone.0231416 — Turns replication prediction into a calibrated predictive distribution rather than a label, and scores it properly — the standard Orion’s confidence calibration should be held to.
  202. 192 Yang, Y., Youyou, W., & Uzzi, B. (2020). Estimating the deep replicability of scientific findings using human and artificial intelligence. Proceedings of the National Academy of Sciences, 117(20), 10762–10768. doi:10.1073/pnas.1909046117 — The headline claim that a model reading only a paper’s text can predict whether its finding will replicate — the strongest published version of the premise Orion depends on, that a stated constraint can be judged from the prose around it.
  203. 193 Guo, Z., Schlichtkrull, M., & Vlachos, A. (2022). A Survey on Automated Fact-Checking. Transactions of the Association for Computational Linguistics, 10, 178–206. doi:10.1162/tacl_a_00454 — The standard map of automated fact-checking — claim detection, evidence retrieval, verdict, justification — against which Orion’s locate, classify, calibrate and present stages can be placed stage for stage.
  204. 194 Wadden, D., Lo, K., Kuehl, B., et al. (2022). SciFact-Open: Towards open-domain scientific claim verification. Findings of the Association for Computational Linguistics: EMNLP 2022, 4719–4734. doi:10.18653/v1/2022.findings-emnlp.347 — The same task reopened over a full corpus rather than a curated candidate pool, and performance falls sharply — the measured cost of moving from a benchmark to the open estate Orion actually sweeps.
  205. 195 Crockett, M. J., Bai, X., Kapoor, S., Messeri, L., & Narayanan, A. (2023). The limitations of machine learning models for predicting scientific replicability. Proceedings of the National Academy of Sciences, 120(33), e2307596120. doi:10.1073/pnas.2307596120 — Cited against this project. The published rebuttal arguing that text-based replicability models are evaluated in ways that overstate them and do not generalise past their training distribution — a live dispute, and the direct counter-evidence to Orion’s ranking premise.
  206. 196 Mottelson, A., & Kontogiorgos, D. (2023). Replicating replicability modeling of psychology papers. Proceedings of the National Academy of Sciences, 120(33), e2309496120. doi:10.1073/pnas.2309496120 — Cited against this project. The second published challenge to text-based replicability prediction, this one an attempt to reproduce the modelling itself; the disagreement is unresolved and both sides are cited here.
  207. 197 Schlichtkrull, M., Ousidhoum, N., & Vlachos, A. (2023). The Intended Uses of Automated Fact-Checking Artefacts: Why, How and Who. Findings of the Association for Computational Linguistics: EMNLP 2023, 8618–8642. doi:10.18653/v1/2023.findings-emnlp.577 — Cited against this project. A survey of what automated fact-checking papers actually claim their systems are for, finding the intended user and the intended decision are usually left unstated — the same gap Orion has to close by naming the analyst who adjudicates.
  208. 198 Wintle, B. C., Smith, E. T., Bush, M., et al. (2023). Predicting and reasoning about replicability using structured groups. Royal Society Open Science, 10(6), 221553. doi:10.1098/rsos.221553 — Structured groups of people, given the paper and a protocol, forecast replication about as well as anything automated — the human comparator Orion’s ranked list is competing with, not the naive baseline.
  209. 199 Youyou, W., Yang, Y., & Uzzi, B. (2023). A discipline-wide investigation of the replicability of Psychology papers over the past two decades. Proceedings of the National Academy of Sciences, 120(6), e2208863120. doi:10.1073/pnas.2208863120 — The same text-based predictor applied across a whole discipline and two decades, which is what drew the rebuttals; the scale claim that has to be weighed against them rather than quoted on its own.
  210. Allocation — choosing what to fund and what to work on nextOrion ranks. These are the literatures on what happens when science chooses, and on the cost of choosing badly.
  211. 200 Merton, R. K. (1968). The Matthew Effect in Science. Science, 159(3810), 56–63. doi:10.1126/science.159.3810.56 — The Matthew effect: recognition accrues to those who already have it — cited against a ranking that, by publishing an order, adds one more mechanism concentrating attention on what is already visible.
  212. 201 Cole, S., Cole, J. R., & Simon, G. A. (1981). Chance and Consensus in Peer Review. Science, 214(4523), 881–886. doi:10.1126/science.7302566 — Putting already-reviewed proposals back through a second independent panel reversed a large share of the decisions — against assuming the human adjudication step at the end of this pipeline is the reliable part.
  213. 202 Keeney, R. L., & Raiffa, H. (1993). Decisions with Multiple Objectives: Preferences and Value Trade-Offs. Cambridge University Press. doi:10.1017/CBO9781139174084 — The theory behind combining several value dimensions into one number, and the conditions under which a multiplicative composite — the form used here — is the correct one rather than a weighted sum.
  214. 203 van Raan, A. F. J. (2004). Sleeping Beauties in science. Scientometrics, 59(3), 467–472. doi:10.1023/B:SCIE.0000018543.82441.f1 — The original identification of the phenomenon of delayed recognition, and the terminology this page inherits.
  215. 204 Evans, J. A. (2008). Electronic Publication and the Narrowing of Science and Scholarship. Science, 321(5887), 395–399. doi:10.1126/science.1150473 — Easier search narrowed rather than broadened what gets cited; the counter-argument to assuming that a searchable record is a fully mined one.
  216. 205 Graves, N., Barnett, A. G., & Clarke, P. (2011). Funding grant proposals for scientific research: retrospective analysis of scores by members of grant review panel. BMJ, 343(sep27 1), d4797–d4797. doi:10.1136/bmj.d4797 — Resampling panel members showed that a large share of decisions near the cut-off would flip on a differently constituted panel — the quantity a ranked candidate list has to publish alongside its order.
  217. 206 Nicholson, J. M., & Ioannidis, J. P. A. (2012). Conform and be funded. Nature, 492(7427), 34–36. doi:10.1038/492034a — Funding tracks conformity to existing high-impact work — the selection pressure that leaves constrained, unfashionable results unrevisited.
  218. 207 Uzzi, B., Mukherjee, S., Stringer, M., & Jones, B. (2013). Atypical Combinations and Scientific Impact. Science, 342(6157), 468–472. doi:10.1126/science.1240474 — Unusual combinations of prior work predict outsized impact — the evidence behind the cross-disciplinary leverage factor.
  219. 208 Bollen, J., Crandall, D., Junk, D., Ding, Y., & Börner, K. (2014). From funding agencies to scientific agency: Collective allocation of science funding as an alternative to peer review. The EMBO Reports, 15(2), 131–133. doi:10.1002/embr.201338068 — A concrete alternative to peer-reviewed allocation; the design space Orion’s ranking would have to compete in.
  220. 209 Foster, J. G., Rzhetsky, A., & Evans, J. A. (2015). Tradition and Innovation in Scientists’ Research Strategies. American Sociological Review, 80(5), 875–908. doi:10.1177/0003122415601618 — Scientists systematically choose safe extensions over risky jumps, and are rewarded for it — the behavioural reason abandoned lines stay abandoned.
  221. 210 Hicks, D., Wouters, P., Waltman, L., de Rijcke, S., & Rafols, I. (2015). Bibliometrics: The Leiden Manifesto for research metrics. Nature, 520(7548), 429–431. doi:10.1038/520429a — Ten principles for the use of quantitative indicators in research assessment, several of which a single composite score of the kind published here plainly violates — cited against this page’s own ranking.
  222. 211 Ke, Q., Ferrara, E., Radicchi, F., & Flammini, A. (2015). Defining and identifying Sleeping Beauties in science. Proceedings of the National Academy of Sciences, 112(24), 7426–7431. doi:10.1073/pnas.1424329112 — Papers whose value is recognised only decades later, defined and counted. The strongest quantitative support for the claim that an old paper can still be undervalued.
  223. 212 Rzhetsky, A., Foster, J. G., Foster, I. T., & Evans, J. A. (2015). Choosing experiments to accelerate collective discovery. Proceedings of the National Academy of Sciences, 112(47), 14569–14574. doi:10.1073/pnas.1509757112 — Treats the choice of the next experiment as an optimisation over the whole field’s effort — the formal statement of what Orion’s ranking is a crude approximation to.
  224. 213 Wilsdon, J. (2015). The Metric Tide: Independent Review of the Role of Metrics in Research Assessment and Management. SAGE Publications Ltd. doi:10.4135/9781473978782 — The independent review of metrics in research assessment, whose central finding is that no indicator is a safe substitute for reading the work — the caution against acting on a rank order alone.
  225. 214 Fang, F. C., & Casadevall, A. (2016). Research Funding: the Case for a Modified Lottery. mBio, 7(2), e00422-16. doi:10.1128/mBio.00422-16 — The argument that near-threshold funding decisions are noise and should be treated as such — the reason a ranking must publish its uncertainty, not just its order.
  226. 215 Hutchins, B. I., Yuan, X., Anderson, J. M., & Santangelo, G. M. (2016). Relative Citation Ratio (RCR): A New Metric That Uses Citation Rates to Measure Influence at the Article Level. PLOS Biology, 14(9), e1002541. doi:10.1371/journal.pbio.1002541 — A field-normalised article-level influence metric built by a funder to make allocation decisions — the state of practice Orion’s composite is an alternative to.
  227. 216 Wang, J., Veugelers, R., & Stephan, P. (2017). Bias against novelty in science: A cautionary tale for users of bibliometric indicators. Research Policy, 46(8), 1416–1436. doi:10.1016/j.respol.2017.06.006 — Novel work is systematically undervalued by the indicators used to allocate funding, and takes longer to pay off — a direct caution about any score, including the composite on this page.
  228. 217 Bol, T., de Vaan, M., & van de Rijt, A. (2018). The Matthew effect in science funding. Proceedings of the National Academy of Sciences, 115(19), 4887–4890. doi:10.1073/pnas.1719557115 — Applicants just above a funding threshold went on to accumulate more than twice the funding of those just below, with no measurable difference in prior quality — the causal measurement of what an arbitrary cut on a ranked list does.
  229. 218 Fortunato, S., Bergstrom, C. T., Börner, K., et al. (2018). Science of science. Science, 359(6379), eaao0185. doi:10.1126/science.aao0185 — The field review for science-of-science methods; the frame in which a system that reads and ranks the literature is a legitimate instrument.
  230. 219 Pier, E. L., Brauer, M., Filut, A., et al. (2018). Low agreement among reviewers evaluating the same NIH grant applications. Proceedings of the National Academy of Sciences, 115(12), 2952–2957. doi:10.1073/pnas.1714379115 — Experienced reviewers scoring the same NIH applications agreed no better than chance — the modern replication of that result, and the ceiling on what any analyst ledger can be validated against.
  231. 220 Azoulay, P., Graff Zivin, J. S., Li, D., & Sampat, B. N. (2019). Public R&D Investments and Private-sector Patenting: Evidence from NIH Funding Rules. The Review of Economic Studies, 86(1), 117–152. doi:10.1093/restud/rdy034 — Causal evidence that public research funding produces private downstream value — the economics under the expected societal impact factor. Crossref records the issued year as 2018 (advance publication); the article appears in the 2019 volume issue.
  232. 221 Gross, K., & Bergstrom, C. T. (2019). Contest models highlight inherent inefficiencies of scientific funding competitions. PLOS Biology, 17(1), e3000065. doi:10.1371/journal.pbio.3000065 — Models the effort scientists burn competing for grants; quantifies what a better allocation mechanism would be worth.
  233. Measuring disruption — and the contested state of that measurementCited with its contestation intact. The headline claim of a disruptiveness decline is under active dispute in the same journal that published it.
  234. 222 Funk, R. J., & Owen-Smith, J. (2017). A Dynamic Network Measure of Technological Change. Management Science, 63(3), 791–817. doi:10.1287/mnsc.2015.2366 — The CD index itself — the measure everything downstream in this group is arguing about.
  235. 223 Bornmann, L., Devarakonda, S., Tekles, A., & Chacko, G. (2020). Are disruption index indicators convergently valid? The comparison of several indicator variants with assessments by peers. Quantitative Science Studies, 1(3), 1242–1259. doi:10.1162/qss_a_00068 — Tests whether disruption indicators agree with expert assessment at all — the validity question underneath the trend question.
  236. 224 Park, M., Leahey, E., & Funk, R. J. (2023). Papers and patents are becoming less disruptive over time. Nature, 613(7942), 138–144. doi:10.1038/s41586-022-05543-x — The headline claim that papers and patents have become less disruptive over time. Cited only together with the two entries below: it is contested, and this page does not present it as settled. Disputed. Part of a live three-way exchange — Park, Leahey & Funk (2023), the Holst et al. Matters Arising, and the authors’ Reply. Read the three together; this page takes no side.
  237. 225 Leibel, C., & Bornmann, L. (2024). What do we know about the disruption index in scientometrics? An overview of the literature. Scientometrics, 129(1), 601–639. doi:10.1007/s11192-023-04873-5 — The literature overview that maps the whole dispute; the entry point for a reader who wants the argument rather than a verdict. Crossref records the issued year as 2023 (advance publication); the article appears in the 2024 volume issue.
  238. 226 Macher, J. T., Rutzer, C., & Weder, R. (2024). Is there a secular decline in disruptive patents? Correcting for measurement bias. Research Policy, 53(5), 104992. doi:10.1016/j.respol.2024.104992 — The same question asked of patents alone, with a correction for measurement bias.
  239. 227 Petersen, A. M., Arroyave, F., & Pammolli, F. (2024). The disruption index is biased by citation inflation. Quantitative Science Studies, 5(4), 936–953. doi:10.1162/qss_a_00333 — The independent objection that the disruption index is biased by citation inflation — a second, separate line of attack on the same measure.
  240. 228 Holst, V., Algaba, A., Tori, F., Wenmackers, S., & Ginis, V. (2026). Dataset artefacts can partially drive the measured decline in disruption. Nature, 656(8127), E7–E13. doi:10.1038/s41586-026-10787-y — The Nature Matters Arising, published 12 August 2026, arguing that dataset artefacts can partially drive the measured decline. Disputed. Part of a live three-way exchange — Park, Leahey & Funk (2023), the Holst et al. Matters Arising, and the authors’ Reply. Read the three together; this page takes no side.
  241. 229 Park, M., Leahey, E., & Funk, R. J. (2026). Reply to: Dataset artefacts can partially drive the measured decline in disruption. Nature, 656(8127), E14–E21. doi:10.1038/s41586-026-10788-x — The authors’ Reply in the same issue. Read the three together or not at all; the exchange is live and this page takes no side in it. Disputed. Part of a live three-way exchange — Park, Leahey & Funk (2023), the Holst et al. Matters Arising, and the authors’ Reply. Read the three together; this page takes no side.
  242. The astronomy worked cases and their data sourcesEvery mission, catalogue and survey named in Findings, plus the three papers taken end to end.
  243. 230 Irwin, J. B. (1952). The Determination of a Light-Time Orbit. The Astrophysical Journal, 116, 211. doi:10.1086/145604 — The light-time effect method the eclipse-timing solutions in worked case two depend on, and the reason a longer baseline is the binding capability. Crossref deposits a start page only.
  244. 231 Reid, I. N., Hawley, S. L., & Gizis, J. E. (1995). The Palomar/MSU Nearby-Star Spectroscopic Survey. I. The Northern M Dwarfs -Bandstrengths and Kinematics. The Astronomical Journal, 110, 1838. doi:10.1086/117655 — The Palomar/MSU survey that supplies the n = 345 sample used in worked case one. Crossref deposits a start page only.
  245. 232 van Leeuwen, F. (2007). Validation of the new Hipparcos reduction. Astronomy & Astrophysics, 474(2), 653–664. doi:10.1051/0004-6361:20078357 — The re-reduced Hipparcos astrometry — the baseline against which the Gaia improvement in worked case three is measured.
  246. 233 Willemsen, P. G., & Eyer, L. (2007). A study of supervised classification of Hipparcos variable stars using PCA and Support Vector Machines. arXiv:0712.2898. arxiv.org/abs/0712.2898 — Worked case three. The paper whose stated constraint — cost of the better stopping criterion — Orion found HELD, then DISSOLVED.
  247. 234 Borucki, W. J., Koch, D., Basri, G., et al. (2010). Kepler Planet-Detection Mission: Introduction and First Results. Science, 327(5968), 977–980. doi:10.1126/science.1185402 — The Kepler mission, source of the eclipse-timing dataset in worked case two.
  248. 235 West, A. A., Morgan, D. P., Bochanski, J. J., et al. (2011). THE SLOAN DIGITAL SKY SURVEY DATA RELEASE 7 SPECTROSCOPIC M DWARF CATALOG. I. DATA. The Astronomical Journal, 141(3), 97. doi:10.1088/0004-6256/141/3/97 — The M dwarf catalogue and activity fractions that worked case one is measured against.
  249. 236 Howell, S. B., Sobeck, C., Haas, M., et al. (2014). The K2 Mission: Characterization and Early Results. Publications of the Astronomical Society of the Pacific, 126(938), 398–408. doi:10.1086/676406 — K2, the extended mission that continued the photometric baseline between Kepler and TESS.
  250. 237 Bailer-Jones, C. A. L. (2015). Estimating Distances from Parallaxes. Publications of the Astronomical Society of the Pacific, 127(956), 994–1009. doi:10.1086/683116 — Why inverting a parallax is not the same as knowing a distance; the statistical care the worked cases had to take before claiming a factor-of-200 improvement.
  251. 238 Ricker, G. R., Winn, J. N., Vanderspek, R., et al. (2015). Transiting Exoplanet Survey Satellite. Journal of Astronomical Telescopes, Instruments, and Systems, 1(1), 014003. doi:10.1117/1.JATIS.1.1.014003 — TESS — the capability that extends the observational baseline from 4.03 to 15.42 years in worked case two. Crossref records the issued year as 2014 (advance publication); the article appears in the 2015 volume issue.
  252. 239 Borkovits, T., Hajdu, T., Sztakovics, J., et al. (2016). A comprehensive study of the Kepler triples via eclipse timing. Monthly Notices of the Royal Astronomical Society, 455(4), 4136–4165. doi:10.1093/mnras/stv2530 — Worked case two. The paper whose stated constraint — insufficient data baseline — Orion tested and found HELD. Crossref records the issued year as 2015 (advance publication); the article appears in the 2016 volume issue.
  253. 240 Gaia Collaboration, Prusti, T., de Bruijne, J. H. J., et al. (2016). The Gaia mission. Astronomy & Astrophysics, 595, A1. doi:10.1051/0004-6361/201629272 — The mission whose parallaxes are the modern capability applied in worked cases one and three.
  254. 241 Jones, D. O., & West, A. A. (2016). A CATALOG OF GALEX ULTRAVIOLET EMISSION FROM SPECTROSCOPICALLY CONFIRMED M DWARFS. The Astrophysical Journal, 817(1), 1. doi:10.3847/0004-637X/817/1/1 — Worked case one. The paper whose stated constraint — distance and temperature uncertainty — Orion tested and found DISSOLVED.
  255. 242 Luri, X., Brown, A. G. A., Sarro, L. M., et al. (2018). Gaia Data Release 2: Using Gaia parallaxes. Astronomy & Astrophysics, 616, A9. doi:10.1051/0004-6361/201832964 — The reference guidance on using Gaia parallaxes correctly, including the zero-point and the traps a naive re-analysis falls into.
  256. 243 Bailer-Jones, C. A. L., Rybizki, J., Fouesneau, M., Demleitner, M., & Andrae, R. (2021). Estimating Distances from Parallaxes. V. Geometric and Photogeometric Distances to 1.47 Billion Stars in Gaia Early Data Release 3. The Astronomical Journal, 161(3), 147. doi:10.3847/1538-3881/abd806 — The distance catalogue actually used, rather than raw inverted parallaxes.
  257. 244 Lindegren, L., Klioner, S. A., Hernández, J., et al. (2021). Gaia Early Data Release 3: The astrometric solution. Astronomy & Astrophysics, 649, A2. doi:10.1051/0004-6361/202039709 — The astrometric solution behind those parallaxes — what the improvement is, and what its residual systematics are.
  258. 245 Gaia Collaboration, Vallenari, A., Brown, A. G. A., et al. (2023). Gaia Data Release 3: Summary of the content and survey properties. Astronomy & Astrophysics, 674, A1. doi:10.1051/0004-6361/202243940 — The specific data release used. Median distance uncertainty falls from 11.0% to 0.052% against the original papers’ distances.
  259. Traffic systems, vehicle routing, fleet and automotive logisticsA stated research interest of this project. Work is aimed here, so the standard surveys and the landmark papers are listed.
  260. 246 Wardrop, J. G. (1952). Some theoretical aspects of road traffic research. Proceedings of the Institution of Civil Engineers, 1(3), 325–362. doi:10.1680/ipeds.1952.11259 — Wardrop’s equilibrium principles — the definition of what a road network settles into, and the basis of all traffic assignment.
  261. 247 Lighthill, M. J., & Whitham, G. B. (1955). On kinematic waves II. A theory of traffic flow on long crowded roads. Proceedings of the Royal Society of London. Series A. Mathematical and Physical Sciences, 229(1178), 317–345. doi:10.1098/rspa.1955.0089 — The LWR kinematic-wave theory — the continuum model underneath essentially all macroscopic traffic modelling.
  262. 248 Richards, P. I. (1956). Shock Waves on the Highway. Operations Research, 4(1), 42–51. doi:10.1287/opre.4.1.42 — The independent derivation of the same shock-wave theory; the other half of the LWR model’s name.
  263. 249 Dantzig, G. B., & Ramser, J. H. (1959). The Truck Dispatching Problem. Management Science, 6(1), 80–91. doi:10.1287/mnsc.6.1.80 — The truck dispatching problem: the paper that created vehicle routing as a field.
  264. 250 Clarke, G., & Wright, J. W. (1964). Scheduling of Vehicles from a Central Depot to a Number of Delivery Points. Operations Research, 12(4), 568–581. doi:10.1287/opre.12.4.568 — The savings algorithm — the first practical heuristic, still the baseline every new method is measured against.
  265. 251 Nagel, K., & Schreckenberg, M. (1992). A cellular automaton model for freeway traffic. Journal de Physique I, 2(12), 2221–2229. doi:10.1051/jp1:1992277 — The cellular-automaton freeway model — the origin of large-scale microscopic traffic simulation.
  266. 252 Daganzo, C. F. (1994). The cell transmission model: A dynamic representation of highway traffic consistent with the hydrodynamic theory. Transportation Research Part B: Methodological, 28(4), 269–287. doi:10.1016/0191-2615(94)90002-7 — The cell transmission model: the discretisation that made kinematic-wave theory computable on a real network.
  267. 253 Treiber, M., Hennecke, A., & Helbing, D. (2000). Congested traffic states in empirical observations and microscopic simulations. Physical Review E, 62(2), 1805–1824. doi:10.1103/PhysRevE.62.1805 — The intelligent driver model, the car-following rule most microsimulation and much automotive control work is built on.
  268. 254 Helbing, D. (2001). Traffic and related self-driven many-particle systems. Reviews of Modern Physics, 73(4), 1067–1141. doi:10.1103/RevModPhys.73.1067 — The comprehensive physics review of traffic as a many-particle system; the reference for congestion phase behaviour.
  269. 255 Golden, B., Raghavan, S., & Wasil, E. (Eds.) (2008). The Vehicle Routing Problem: Latest Advances and New Challenges. Operations Research/Computer Science Interfaces. Springer US. doi:10.1007/978-0-387-77778-8 — The companion collection covering the variants and open challenges the standard volume compresses. Crossref deposits these names in the editor role rather than the author role; they are shown here as editors.
  270. 256 Laporte, G. (2009). Fifty Years of Vehicle Routing. Transportation Science, 43(4), 408–416. doi:10.1287/trsc.1090.0301 — Fifty years of the field in one review; the orientation piece for anyone entering this area.
  271. 257 Bektaş, T., & Laporte, G. (2011). The Pollution-Routing Problem. Transportation Research Part B: Methodological, 45(8), 1232–1250. doi:10.1016/j.trb.2011.02.004 — The pollution-routing problem — routing where the objective includes emissions and speed, not only distance.
  272. 258 Erdoğan, S., & Miller-Hooks, E. (2012). A Green Vehicle Routing Problem. Transportation Research Part E: Logistics and Transportation Review, 48(1), 100–114. doi:10.1016/j.tre.2011.08.001 — The green VRP: routing with refuelling and range constraints, the formulation alternative-fuel fleets inherit.
  273. 259 Pillac, V., Gendreau, M., Guéret, C., & Medaglia, A. L. (2013). A review of dynamic vehicle routing problems. European Journal of Operational Research, 225(1), 1–11. doi:10.1016/j.ejor.2012.08.015 — Dynamic vehicle routing: what changes when the problem is re-solved as information arrives, which is the realistic fleet case.
  274. 260 Vidal, T., Crainic, T. G., Gendreau, M., & Prins, C. (2013). Heuristics for multi-attribute vehicle routing problems: A survey and synthesis. European Journal of Operational Research, 231(1), 1–21. doi:10.1016/j.ejor.2013.02.053 — The survey of heuristics for multi-attribute routing — what actually solves real fleet problems at size.
  275. 261 Toth, P., & Vigo, D. (Eds.) (2014). Vehicle Routing: Problems, Methods, and Applications, Second Edition. Society for Industrial and Applied Mathematics. doi:10.1137/1.9781611973594 — The standard reference volume on vehicle routing problems, methods and applications. Crossref deposits these names in the editor role rather than the author role; they are shown here as editors.
  276. 262 Vlahogianni, E. I., Karlaftis, M. G., & Golias, J. C. (2014). Short-term traffic forecasting: Where we are and where we’re going. Transportation Research Part C: Emerging Technologies, 43, 3–19. doi:10.1016/j.trc.2014.01.005 — Short-term traffic forecasting: what the field had established and what it had not, which is where a constraint-mining pass would look for stated limits.
  277. 263 Braekers, K., Ramaekers, K., & Van Nieuwenhuyse, I. (2016). The vehicle routing problem: State of the art classification and review. Computers & Industrial Engineering, 99, 300–313. doi:10.1016/j.cie.2015.12.007 — The most-used modern classification of VRP variants; the taxonomy to place a new routing problem in.
  278. 264 Pelletier, S., Jabali, O., & Laporte, G. (2016). 50th Anniversary Invited Article—Goods Distribution with Electric Vehicles: Review and Research Perspectives. Transportation Science, 50(1), 3–22. doi:10.1287/trsc.2015.0646 — Goods distribution with electric vehicles: the review that sets out what electrification does to fleet planning.
  279. 265 Cattaruzza, D., Absi, N., Feillet, D., & González-Feliu, J. (2017). Vehicle routing problems for city logistics. EURO Journal on Transportation and Logistics, 6(1), 51–79. doi:10.1007/s13676-014-0074-0 — City logistics — routing under the constraints that actually bind in urban delivery.
  280. Medical and pharmaceutical supply chain, cold chain, hospital and humanitarian health logisticsThe second stated research interest. Same standard: the surveys a domain expert would expect, and the papers they descend from.
  281. 266 Shah, N. (2004). Pharmaceutical supply chains: key issues and strategies for optimisation. Computers & Chemical Engineering, 28(6-7), 929–941. doi:10.1016/j.compchemeng.2003.09.022 — The framing survey of pharmaceutical supply chains and where the optimisation opportunities sit.
  282. 267 Altay, N., & Green, W. G. (2006). OR/MS research in disaster operations management. European Journal of Operational Research, 175(1), 475–493. doi:10.1016/j.ejor.2005.05.016 — The survey of operational-research work in disaster operations; the method inventory for the humanitarian case.
  283. 268 Van Wassenhove, L. N. (2006). Humanitarian aid logistics: supply chain management in high gear. Journal of the Operational Research Society, 57(5), 475–489. doi:10.1057/palgrave.jors.2602125 — The paper that established humanitarian logistics as a discipline distinct from commercial supply chain management.
  284. 269 Kovács, G., & Spens, K. M. (2007). Humanitarian logistics in disaster relief operations. International Journal of Physical Distribution & Logistics Management, 37(2), 99–114. doi:10.1108/09600030710734820 — Humanitarian logistics in disaster relief — the operational description that most later modelling work is calibrated against.
  285. 270 Beamon, B. M., & Balcik, B. (2008). Performance measurement in humanitarian relief chains. International Journal of Public Sector Management, 21(1), 4–25. doi:10.1108/09513550810846087 — How to measure performance when the objective is not cost — the measurement problem in humanitarian health logistics.
  286. 271 Balcik, B., Beamon, B. M., Krejci, C. C., Muramatsu, K. M., & Ramirez, M. (2010). Coordination in humanitarian relief chains: Practices, challenges and opportunities. International Journal of Production Economics, 126(1), 22–34. doi:10.1016/j.ijpe.2009.09.008 — Coordination failure between relief actors, which is the constraint that dominates the routing and inventory ones.
  287. 272 Mackey, T. K., & Liang, B. A. (2011). The global counterfeit drug trade: Patient safety and public health risks. Journal of Pharmaceutical Sciences, 100(11), 4571–4579. doi:10.1002/jps.22679 — Counterfeit penetration of the drug supply — why traceability and chain-of-custody are supply chain design constraints, not compliance overhead.
  288. 273 Rais, A., & Viana, A. (2011). Operations Research in Healthcare: a survey. International Transactions in Operational Research, 18(1), 1–31. doi:10.1111/j.1475-3995.2010.00767.x — The broad survey of operational research in healthcare; the frame that connects facility, inventory and routing decisions. Crossref records the issued year as 2010 (advance publication); the article appears in the 2011 volume issue.
  289. 274 Beliën, J., & Forcé, H. (2012). Supply chain management of blood products: A literature review. European Journal of Operational Research, 217(1), 1–16. doi:10.1016/j.ejor.2011.05.026 — Blood products: a perishable, uncertain-supply, uncertain-demand chain; the hardest standard case in health logistics.
  290. 275 Privett, N., & Gonsalvez, D. (2014). The top ten global health supply chain issues: Perspectives from the field. Operations Research for Health Care, 3(4), 226–230. doi:10.1016/j.orhc.2014.09.002 — The practitioner-derived list of the ten issues that actually break global health supply chains — the reality check on model-side work.
  291. 276 Yadav, P. (2015). Health Product Supply Chains in Developing Countries: Diagnosis of the Root Causes of Underperformance and an Agenda for Reform. Health Systems & Reform, 1(2), 142–154. doi:10.4161/23288604.2014.968005 — Diagnoses why health product supply chains in developing countries underperform; the problem statement humanitarian health logistics work responds to.
  292. 277 Lemmens, S., Decouttere, C., Vandaele, N., & Bernuzzi, M. (2016). A review of integrated supply chain network design models: Key issues for vaccine supply chains. Chemical Engineering Research and Design, 109, 366–384. doi:10.1016/j.cherd.2016.02.015 — Network design for vaccine supply chains — the strategic layer above routing and inventory.
  293. 278 Ahmadi-Javid, A., Seyedi, P., & Syam, S. S. (2017). A survey of healthcare facility location. Computers & Operations Research, 79, 223–263. doi:10.1016/j.cor.2016.05.018 — Healthcare facility location — the strategic decision that fixes what any downstream routing can achieve.
  294. 279 Fikar, C., & Hirsch, P. (2017). Home health care routing and scheduling: A review. Computers & Operations Research, 77, 86–95. doi:10.1016/j.cor.2016.07.019 — Home health care routing and scheduling — the point where vehicle routing and health logistics, this project’s two stated interests, are the same problem.
  295. 280 Lee, B. Y., & Haidari, L. A. (2017). The importance of vaccine supply chains to everyone in the vaccine world. Vaccine, 35(35), 4475–4479. doi:10.1016/j.vaccine.2017.05.096 — Why the supply chain, not the vaccine, is often the binding constraint on immunisation coverage.
  296. 281 Settanni, E., Harrington, T. S., & Srai, J. S. (2017). Pharmaceutical supply chain models: A synthesis from a systems view of operations research. Operations Research Perspectives, 4, 74–95. doi:10.1016/j.orp.2017.05.002 — A systems-view synthesis of pharmaceutical supply chain models — the map of which operational-research formulations have actually been applied.
  297. 282 Volland, J., Fügener, A., Schoenfelder, J., & Brunner, J. O. (2017). Material logistics in hospitals: A literature review. Omega, 69, 82–101. doi:10.1016/j.omega.2016.08.004 — The literature review of material logistics inside hospitals, where the supply chain terminates.
  298. 283 Comes, T., Bergtora Sandvik, K., & Van de Walle, B. (2018). Cold chains, interrupted: The use of technology and information for decisions that keep humanitarian vaccines cool. Journal of Humanitarian Logistics and Supply Chain Management, 8(1), 49–69. doi:10.1108/JHLSCM-03-2017-0006 — Cold chain failure in the field, and the information gap that causes it — the direct source on humanitarian cold chain.
  299. 284 Duijzer, L. E., van Jaarsveld, W., & Dekker, R. (2018). Literature review: The vaccine supply chain. European Journal of Operational Research, 268(1), 174–192. doi:10.1016/j.ejor.2018.01.015 — The reference review of the vaccine supply chain, end to end from production to administration.
  300. 285 Ahmadi, E., Masel, D. T., Metcalf, A. Y., & Schuller, K. (2019). Inventory management of surgical supplies and sterile instruments in hospitals: a literature review. Health Systems, 8(2), 134–151. doi:10.1080/20476965.2018.1496875 — Surgical supply and sterile instrument inventory — the highest-value, most constrained inventory problem in a hospital. Crossref records the issued year as 2018 (advance publication); the article appears in the 2019 volume issue.
  301. Retrieval, ranking and representationThe methods the locate-and-rank stages are built on.
  302. 286 Sparck Jones, K. (1972). A statistical interpretation of term specificity and its application in retrieval. Journal of Documentation, 28(1), 11–21. doi:10.1108/eb026526 — Inverse document frequency — the reason a trigger phrase is gated on document frequency before it enters the vocabulary.
  303. 287 Viola, P., & Jones, M. (2001). Rapid object detection using a boosted cascade of simple features. Proceedings of the 2001 IEEE Computer Society Conference on Computer Vision and Pattern Recognition. CVPR 2001, 1, I–511–I–518. doi:10.1109/CVPR.2001.990517 — The original attentional cascade: reject most candidates cheaply, spend effort only on survivors. The architectural ancestor of the funnel on this page. Crossref carries no issued date for this deposit; the year is taken from the proceedings title in the same record.
  304. 288 Robertson, S., & Zaragoza, H. (2009). The Probabilistic Relevance Framework: BM25 and Beyond. Foundations and Trends® in Information Retrieval, 4(1-2), 1–174. doi:10.1561/1500000019 — The probabilistic relevance framework and BM25; the scoring model behind lexical retrieval over the estate.
  305. 289 Wang, L., Lin, J., & Metzler, D. (2011). A cascade ranking model for efficient ranked retrieval. Proceedings of the 34th international ACM SIGIR conference on Research and development in Information Retrieval, 105–114. doi:10.1145/2009916.2009934 — The formal cascade ranking model — why staging a cheap filter before an expensive one is the right shape for this problem.
  306. 290 Mikolov, T., Chen, K., Corrado, G., & Dean, J. (2013). Efficient Estimation of Word Representations in Vector Space. arXiv:1301.3781. doi:10.48550/arXiv.1301.3781 — Word vectors — the result that made distributional similarity cheap enough to expand a phrase list by neighbourhood.
  307. 291 Vaswani, A., Shazeer, N., Parmar, N., et al. (2017). Attention Is All You Need. arXiv:1706.03762. doi:10.48550/arXiv.1706.03762 — The transformer architecture every model in this pipeline, from the embedder to the judge, is an instance of.
  308. 292 Peters, M., Neumann, M., Iyyer, M., et al. (2018). Deep Contextualized Word Representations. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), 2227–2237. doi:10.18653/v1/n18-1202 — Contextual word representations before the transformer — the demonstration that a word’s reading depends on its neighbours, which is the whole reason Orion classifies limitation sentences in context rather than by keyword.
  309. 293 Beltagy, I., Lo, K., & Cohan, A. (2019). SciBERT: A Pretrained Language Model for Scientific Text. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 3613–3618. doi:10.18653/v1/D19-1371 — A language model pretrained on scientific text rather than general text; the correct starting point for reading papers.
  310. 294 Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the 2019 Conference of the North, 4171–4186. doi:10.18653/v1/N19-1423 — Contextual sentence representations — the model class that makes a span classifier possible at all. Crossref carries this deposit without a title; the title and pagination were read from the ACL Anthology record.
  311. 295 Dodge, J., Gururangan, S., Card, D., Schwartz, R., & Smith, N. A. (2019). Show Your Work: Improved Reporting of Experimental Results. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 2185–2194. doi:10.18653/v1/d19-1224 — Cited against this project. Shows that reported results are a function of the search budget spent on them, so a comparison that omits the budget is uninterpretable — the reporting standard Orion’s own numbers have to be held to.
  312. 296 Gorman, K., & Bedrick, S. (2019). We Need to Talk about Standard Splits. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2786–2791. doi:10.18653/v1/p19-1267 — Cited against this project. Demonstrates that system rankings on standard splits do not survive random re-splitting — the warning that any Orion comparison between limitation classifiers may be an artefact of one partition.
  313. 297 Liu, Y., Ott, M., Goyal, N., et al. (2019). RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv. doi:10.48550/arXiv.1907.11692 — The controlled re-training that showed BERT was badly under-trained, and that much of what had been credited to architecture was really data and schedule — the caution behind attributing any Orion result to model choice. arXiv:1907.11692; the DOI is registered with DataCite.
  314. 298 Nogueira, R., & Cho, K. (2019). Passage Re-ranking with BERT. arXiv:1901.04085. doi:10.48550/arXiv.1901.04085 — Re-ranking a cheap candidate list with an expensive model; the cascade Orion’s locate-then-judge sequence implements.
  315. 299 Raffel, C., Shazeer, N., Roberts, A., et al. (2019). Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. arXiv. doi:10.48550/arXiv.1910.10683 — The text-to-text formulation that turned classification into generation — the standing alternative to Orion’s discriminative classifier, and the design it has to be measured against. arXiv:1910.10683; the DOI is registered with DataCite.
  316. 300 Reimers, N., & Gurevych, I. (2019). Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 3980–3990. doi:10.18653/v1/D19-1410 — Sentence embeddings that can be compared directly — the mechanism for exemplar-to-candidate phrase expansion.
  317. 301 Strubell, E., Ganesh, A., & McCallum, A. (2019). Energy and Policy Considerations for Deep Learning in NLP. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 3645–3650. doi:10.18653/v1/p19-1355 — Cited against this project. Puts a number on the energy and the money a single training run consumes — the direct argument that re-training an Orion classifier repeatedly over a large estate is not free.
  318. 302 Beltagy, I., Peters, M. E., & Cohan, A. (2020). Longformer: The Long-Document Transformer. arXiv. doi:10.48550/arXiv.2004.05150 — Attention that scales to long documents — the direct answer to the fact that a stated limitation usually sits near the end of a full paper, far past a 512-token window. arXiv:2004.05150; the DOI is registered with DataCite.
  319. 303 Brown, T. B., Mann, B., Ryder, N., et al. (2020). Language Models are Few-Shot Learners. arXiv. doi:10.48550/arXiv.2005.14165 — Few-shot prompting offered as a substitute for supervised training — the route Orion would be forced onto if labelled limitation sentences could not be obtained in quantity. arXiv:2005.14165; the DOI is registered with DataCite.
  320. 304 Cohan, A., Feldman, S., Beltagy, I., Downey, D., & Weld, D. (2020). SPECTER: Document-level Representation Learning using Citation-informed Transformers. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2270–2282. doi:10.18653/v1/2020.acl-main.207 — Document-level scientific embeddings informed by the citation graph; the representation for whole-paper similarity and de-duplication.
  321. 305 Desai, S., & Durrett, G. (2020). Calibration of Pre-trained Transformers. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 295–302. doi:10.18653/v1/2020.emnlp-main.21 — Cited against this project. Finds pre-trained transformers reasonably calibrated in domain but badly overconfident out of it — exactly the regime Orion enters when it sweeps disciplines its classifier was not tuned on.
  322. 306 Gururangan, S., Marasović, A., Swayamdipta, S., et al. (2020). Don’t Stop Pretraining: Adapt Language Models to Domains and Tasks. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 8342–8360. doi:10.18653/v1/2020.acl-main.740 — The systematic measurement of what continued pre-training on in-domain and in-task text actually buys — evidence for the claim that a general encoder should be adapted to the scientific estate before it is asked to classify anything in it.
  323. 307 Joshi, M., Chen, D., Liu, Y., Weld, D. S., Zettlemoyer, L., & Levy, O. (2020). SpanBERT: Improving Pre-training by Representing and Predicting Spans. Transactions of the Association for Computational Linguistics, 8, 64–77. doi:10.1162/tacl_a_00300 — Pre-training that predicts whole spans rather than single tokens — the representation shape Orion needs, because a stated constraint is a span of a sentence and not a word.
  324. 308 Karpukhin, V., Oguz, B., Min, S., et al. (2020). Dense Passage Retrieval for Open-Domain Question Answering. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 6769–6781. doi:10.18653/v1/2020.emnlp-main.550 — Dense passage retrieval — the learned alternative to a phrase sweep, and the obvious upgrade path for the Locate stage.
  325. 309 Lewis, P., Perez, E., Piktus, A., et al. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. arXiv:2005.11401. doi:10.48550/arXiv.2005.11401 — Retrieval-augmented generation — the pattern of grounding a model’s judgement in retrieved text rather than its own parameters, which is what keeps a span decision attributable to the document.
  326. 310 Malkov, Y. A., & Yashunin, D. A. (2020). Efficient and Robust Approximate Nearest Neighbor Search Using Hierarchical Navigable Small World Graphs. IEEE Transactions on Pattern Analysis and Machine Intelligence, 42(4), 824–836. doi:10.1109/TPAMI.2018.2889473 — HNSW — the graph index that makes nearest-neighbour search over a corpus-scale embedding set tractable.
  327. 311 Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the Dangers of Stochastic Parrots. Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, 610–623. doi:10.1145/3442188.3445922 — Cited against this project. The catalogue of costs in scaling language models — environmental, financial and evaluative — set against any plan to throw a larger encoder at a corpus sweep. The ACM deposit records the title without its subtitle. The registrant deposits the main title only; the subtitle after the colon is not in the record, so it is not printed here.
  328. 312 Johnson, J., Douze, M., & Jegou, H. (2021). Billion-Scale Similarity Search with GPUs. IEEE Transactions on Big Data, 7(3), 535–547. doi:10.1109/TBDATA.2019.2921572 — Billion-scale similarity search on GPUs; the practical ceiling on how large the searchable estate can get.
  329. 313 Peng, B., Chersoni, E., Hsu, Y. Y., & Huang, C. R. (2021). Is Domain Adaptation Worth Your Investment? Comparing BERT and FinBERT on Financial Tasks. Proceedings of the Third Workshop on Economics and Natural Language Processing, 37–44. doi:10.18653/v1/2021.econlp-1.5 — Cited against this project. A controlled comparison in which the domain-adapted encoder beats the general one by only a slim margin on domain tasks — the cost case against pre-training a bespoke encoder for Orion’s corpus.
  330. 314 Wei, J., Bosma, M., Zhao, V. Y., et al. (2021). Finetuned Language Models Are Zero-Shot Learners. arXiv. doi:10.48550/arXiv.2109.01652 — Instruction tuning, which made zero-shot task following practical — the reason a limitation classifier can now be specified in a prompt rather than trained, and the baseline a trained Orion classifier has to beat. arXiv:2109.01652; the DOI is registered with DataCite.
  331. 315 Steck, H., Ekanadham, C., & Kallus, N. (2024). Is Cosine-Similarity of Embeddings Really About Similarity?. Companion Proceedings of the ACM Web Conference 2024, 887–890. doi:10.1145/3589335.3651526 — Cited against this project. Shows that cosine similarity between learned embeddings can be close to arbitrary, set by the regularisation rather than by meaning — which undercuts any Orion ranking that treats embedding proximity as a proxy for relevance.
  332. Language models as judges — and the limits of thatStages 3 and 4 are closed-set model judgements. These are the sources for what that buys and what it costs.
  333. 316 Lipton, Z. C. (2018). The Mythos of Model Interpretability: In machine learning, the concept of interpretability is both important and slippery. Queue, 16(3), 31–57. doi:10.1145/3236386.3241340 — What interpretability does and does not mean — the caution against reading a rule-derived factor score as an explanation.
  334. 317 Mitchell, M., Wu, S., Zaldivar, A., et al. (2019). Model Cards for Model Reporting. Proceedings of the Conference on Fairness, Accountability, and Transparency, 220–229. doi:10.1145/3287560.3287596 — Model cards — the disclosure standard for stating what a deployed model was evaluated on and where it is known to fail.
  335. 318 Hendrycks, D., Burns, C., Basart, S., et al. (2020). Measuring Massive Multitask Language Understanding. arXiv:2009.03300. doi:10.48550/arXiv.2009.03300 — The closed-set multiple-choice evaluation format that the judge stages borrow — forcing a choice from a fixed set rather than accepting free text.
  336. 319 Gebru, T., Morgenstern, J., Vecchione, B., et al. (2021). Datasheets for datasets. Communications of the ACM, 64(12), 86–92. doi:10.1145/3458723 — Datasheets for datasets; the standard behind publishing the estate manifest, the exclusions and the retrieval failures rather than only the survivors.
  337. 320 Alkaissi, H., & McFarlane, S. I. (2023). Artificial Hallucinations in ChatGPT: Implications in Scientific Writing. Cureus. doi:10.7759/cureus.35179 — Cited against this project. The early case report that named the failure: a model producing fluent scientific prose supported by references that do not exist.
  338. 321 Bhattacharyya, M., Miller, V. M., Bhattacharyya, D., & Miller, L. E. (2023). High Rates of Fabricated and Inaccurate References in ChatGPT-Generated Medical Content. Cureus. doi:10.7759/cureus.39238 — Cited against this project. Found most references in model-generated medical content to be fabricated or inaccurate — the same failure again, in the domain where a wrong citation does the most damage.
  339. 322 Gao, C. A., Howard, F. M., Markov, N. S., et al. (2023). Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewers. npj Digital Medicine, 6(1), 75. doi:10.1038/s41746-023-00819-6 — Model-written abstracts that blinded human reviewers could not reliably separate from real ones — contested ground for any pipeline that asks a reader to judge machine-assembled scientific text on its face.
  340. 323 Gilardi, F., Alizadeh, M., & Kubli, M. (2023). ChatGPT outperforms crowd workers for text-annotation tasks. Proceedings of the National Academy of Sciences, 120(30), e2305016120. doi:10.1073/pnas.2305016120 — Evidence for a claim here: across several annotation tasks a model matched or beat crowd workers at a fraction of the cost — the empirical case for model labelling where Orion would otherwise need paid annotators.
  341. 324 Ji, Z., Lee, N., Frieske, R., et al. (2023). Survey of Hallucination in Natural Language Generation. ACM Computing Surveys, 55(12), 1–38. doi:10.1145/3571730 — The survey of generation that is fluent and wrong; the reason every quoted span on this page carries offsets read from the document rather than from a model.
  342. 325 Liu, Y., Iter, D., Xu, Y., Wang, S., Xu, R., & Zhu, C. (2023). G-Eval: NLG Evaluation using Gpt-4 with Better Human Alignment. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2511–2522. doi:10.18653/v1/2023.emnlp-main.153 — The chain-of-thought scoring framework that made model-graded evaluation a named method — the technique Orion would be using if a model, rather than an analyst, scored candidate constraints. The Crossref deposit capitalises the model name as &ldquo;Gpt-4&rdquo;.
  343. 326 Walters, W. H., & Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports, 13(1), 14045. doi:10.1038/s41598-023-41032-5 — Cited against this project. A measured rate of fabricated and erroneous references in model-generated bibliographies — the reason nothing on this page was written from a model’s memory.
  344. 327 Zheng, L., Chiang, W.-L., Sheng, Y., et al. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. arXiv:2306.05685. doi:10.48550/arXiv.2306.05685 — The reference study on using a model as a judge: where it agrees with human raters and where it does not. Stages 3 and 4 of this pipeline are exactly this, so its failure modes are this page’s failure modes.
  345. 328 Koo, R., Lee, M., Raheja, V., Park, J. I., Kim, Z. M., & Kang, D. (2024). Benchmarking Cognitive Biases in Large Language Models as Evaluators. Findings of the Association for Computational Linguistics ACL 2024, 517–545. doi:10.18653/v1/2024.findings-acl.29 — Cited against this project. A benchmark that measures cognitive biases in model evaluators rather than asserting them — the systematic evidence that model-graded scoring is not a neutral instrument.
  346. 329 Panickssery, A., Bowman, S., & Feng, S. (2024). LLM Evaluators Recognize and Favor Their Own Generations. Advances in Neural Information Processing Systems 37, 68772–68802. doi:10.52202/079017-2197 — Cited against this project. Model evaluators recognise and prefer their own output — self-preference bias, which makes any pipeline that both generates and scores its own candidates circular.
  347. 330 Wang, P., Li, L., Chen, L., et al. (2024). Large Language Models are not Fair Evaluators. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 9440–9450. doi:10.18653/v1/2024.acl-long.511 — Cited against this project. Establishes that a model judge’s verdict can flip when the two candidates swap places — positional bias, and a direct reason not to let a model adjudicate Orion’s ranked list.
  348. 331 Ziems, C., Held, W., Shaikh, O., Chen, J., Zhang, Z., & Yang, D. (2024). Can Large Language Models Transform Computational Social Science?. Computational Linguistics, 50(1), 237–291. doi:10.1162/coli_a_00502 — Measures where large language models can and cannot replace human annotation on social-scientific labelling — the calibration for how far the closed-set judge’s verdicts can be trusted unaided.
  349. 332 Shi, L., Ma, C., Liang, W., Diao, X., Ma, W., & Vosoughi, S. (2025). Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge. Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics, 292–314. doi:10.18653/v1/2025.ijcnlp-long.18 — Cited against this project. A systematic study of position bias across model judges and benchmarks — it quantifies how much of a model-graded verdict is an artefact of ordering alone.
  350. 333 Szymanski, A., Ziems, N., Eicher-Miller, H. A., Li, T. J. J., Jiang, M., & Metoyer, R. A. (2025). Limitations of the LLM-as-a-Judge Approach for Evaluating LLM Outputs in Expert Knowledge Tasks. Proceedings of the 30th International Conference on Intelligent User Interfaces, 952–966. doi:10.1145/3708359.3712091 — Cited against this project. Compares model judges with domain experts on expert-knowledge tasks and finds the model verdicts do not line up with expert judgement — the closest published test of the adjudication Orion hands to a human.
  351. Screening and triage at corpus scaleThe funnel that turns an estate manifest into a candidate set is a screening problem with its own literature.
  352. 334 Voorhees, E. M. (2000). Variations in relevance judgments and the measurement of retrieval effectiveness. Information Processing & Management, 36(5), 697–716. doi:10.1016/s0306-4573(00)00010-8 — Different assessors disagree substantially about which documents are relevant, and absolute effectiveness scores move with the assessor — against treating one analyst’s adjudication of a candidate list as ground truth.
  353. 335 Cohen, A. M., Hersh, W. R., Peterson, K., & Yen, P. Y. (2006). Reducing Workload in Systematic Review Preparation Using Automated Citation Classification. Journal of the American Medical Informatics Association, 13(2), 206–219. doi:10.1197/jamia.m1929 — The founding demonstration that automated citation classification can cut the human reading load in a systematic review — the shape of Orion’s output, a ranked list that saves an analyst time rather than replacing the analyst.
  354. 336 Wallace, B. C., Trikalinos, T. A., Lau, J., Brodley, C., & Schmid, C. H. (2010). Semi-automated screening of biomedical citations for systematic reviews. BMC Bioinformatics, 11(1), 55. doi:10.1186/1471-2105-11-55 — Semi-automated citation screening — the established design for a machine that triages and a human who adjudicates, which is the analyst ledger on this page.
  355. 337 Wallace, B. C., Small, K., Brodley, C. E., & Trikalinos, T. A. (2010). Active learning for biomedical citation screening. Proceedings of the 16th ACM SIGKDD international conference on Knowledge discovery and data mining, 173–182. doi:10.1145/1835804.1835829 — Active learning applied to citation screening, where the machine chooses which abstract the human reads next — the query strategy an analyst ledger implies.
  356. 338 Cormack, G. V., & Grossman, M. R. (2014). Evaluation of machine-learning protocols for technology-assisted review in electronic discovery. Proceedings of the 37th international ACM SIGIR conference on Research & development in information retrieval, 153–162. doi:10.1145/2600428.2609601 — The controlled comparison of continuous active learning against simple passive learning for technology-assisted review — the protocol the triage loop here is a variant of.
  357. 339 Marshall, I. J., Kuiper, J., & Wallace, B. C. (2015). RobotReviewer: evaluation of a system for automatically assessing bias in clinical trials. Journal of the American Medical Informatics Association, 23(1), 193–201. doi:10.1093/jamia/ocv044 — A deployed system that reads trial reports, marks risk of bias and shows the sentences behind each mark, evaluated as a whole — the closest working analogue to Orion’s present-the-evidence-to-an-analyst step.
  358. 340 O’Mara-Eves, A., Thomas, J., McNaught, J., Miwa, M., & Ananiadou, S. (2015). Using text mining for study identification in systematic reviews: a systematic review of current approaches. Systematic Reviews, 4(1), 5. doi:10.1186/2046-4053-4-5 — The systematic review of text mining for study identification: what recall is achievable, at what workload saving, and under what evaluation.
  359. 341 Cormack, G. V., & Grossman, M. R. (2016). Engineering Quality and Reliability in Technology-Assisted Review. Proceedings of the 39th International ACM SIGIR conference on Research and Development in Information Retrieval, 75–84. doi:10.1145/2911451.2911510 — Stopping rules for technology-assisted review, and the recall guarantee a reviewer can actually be given — the missing piece in any ranked candidate list a human has to stop working down.
  360. 342 Howard, B. E., Phillips, J., Miller, K., et al. (2016). SWIFT-Review: a text-mining workbench for systematic review. Systematic Reviews, 5(1), 87. doi:10.1186/s13643-016-0263-z — A deployed text-mining workbench for prioritising and triaging a literature — the nearest working system to the analyst-facing half of this pipeline.
  361. 343 Gates, A., Johnson, C., & Hartling, L. (2018). Technology-assisted title and abstract screening for systematic reviews: a retrospective evaluation of the Abstrackr machine learning tool. Systematic Reviews, 7(1), 45. doi:10.1186/s13643-018-0707-8 — A retrospective evaluation of a widely used screening tool across real reviews found recall varying unpredictably from one review to the next — cited against the assumption that a triage classifier scoring well on one corpus will hold up on the next.
  362. 344 Gates, A., Guitard, S., Pillay, J., et al. (2019). Performance and usability of machine learning for screening in systematic reviews: a comparative evaluation of three tools. Systematic Reviews, 8(1), 278. doi:10.1186/s13643-019-1222-2 — Three screening tools run head to head on the same reviews, with recall and workload saving that differ sharply between them and between reviews — against quoting a single triage performance figure as if it were a property of the method.
  363. 345 Marshall, I. J., & Wallace, B. C. (2019). Toward systematic review automation: a practical guide to using machine learning tools in research synthesis. Systematic Reviews, 8(1), 163. doi:10.1186/s13643-019-1074-9 — The practical guide to which review-automation tools are ready for use and which are not — the state of the art the triage half of this pipeline has to be measured against.
  364. 346 O’Connor, A. M., Tsafnat, G., Thomas, J., Glasziou, P., Gilbert, S. B., & Hutton, B. (2019). A question of trust: can we build an evidence base to gain trust in systematic review automation technologies?. Systematic Reviews, 8(1), 143. doi:10.1186/s13643-019-1062-0 — Argues the evidence base for trusting review-automation tools does not yet exist and sets out what would have to be measured first — a standard of proof this pipeline has not met either.
  365. 347 Callaghan, M. W., & Müller-Hansen, F. (2020). Statistical stopping criteria for automated screening in systematic reviews. Systematic Reviews, 9(1), 273. doi:10.1186/s13643-020-01521-4 — A statistical stopping criterion with an explicit recall guarantee for machine-assisted screening — how to say when enough of a ranked candidate list has been read.
  366. The statistics used on this siteEvery interval, effect size and correction the page computes, traced to its source.
  367. 348 Pearson, K. (1900). On the criterion that a given system of deviations from the probable in the case of a correlated system of variables is such that it can be reasonably supposed to have arisen from random sampling. The London, Edinburgh, and Dublin Philosophical Magazine and Journal of Science, 50(302), 157–175. doi:10.1080/14786440009463897 — The chi-squared statistic that any association measure over the contingency tables on this page is derived from.
  368. 349 Wilson, E. B. (1927). Probable Inference, the Law of Succession, and Statistical Inference. Journal of the American Statistical Association, 22(158), 209–212. doi:10.1080/01621459.1927.10502953 — The score interval used for every proportion on this page.
  369. 350 CLOPPER, C. J., & PEARSON, E. S. (1934). THE USE OF CONFIDENCE OR FIDUCIAL LIMITS ILLUSTRATED IN THE CASE OF THE BINOMIAL. Biometrika, 26(4), 404–413. doi:10.1093/biomet/26.4.404 — The exact binomial interval — the conservative baseline the score interval used on this page is chosen over, and the reference point for what “exact” costs in width.
  370. 351 Cramér, H. (1946). Mathematical Methods of Statistics (PMS-9). Princeton University Press. doi:10.1515/9781400883868 — Cramér’s V — the normalisation of chi-squared into an effect size comparable across tables of different size, which is what makes cross-cohort association numbers legible.
  371. 352 Shannon, C. E. (1948). A Mathematical Theory of Communication. Bell System Technical Journal, 27(3), 379–423. doi:10.1002/j.1538-7305.1948.tb01338.x — The definition of entropy and mutual information that every term-weighting and feature-selection step in this pipeline is a special case of.
  372. 353 BRIER, G. W. (1950). VERIFICATION OF FORECASTS EXPRESSED IN TERMS OF PROBABILITY. Monthly Weather Review, 78(1), 1–3. doi:10.1175/1520-0493(1950)078<0001:vofeit>2.0.co;2 — The Brier score — the proper scoring rule the calibrated confidences on this page are graded by.
  373. 354 Ayer, M., Brunk, H. D., Ewing, G. M., Reid, W. T., & Silverman, E. (1955). An Empirical Distribution Function for Sampling with Incomplete Information. The Annals of Mathematical Statistics, 26(4), 641–647. doi:10.1214/aoms/1177728423 — The pool-adjacent-violators result underneath isotonic regression — the estimator whose monotone fit the calibration stage relies on.
  374. 355 Kitagawa, E. M. (1955). Components of a Difference Between Two Rates. Journal of the American Statistical Association, 50(272), 1168–1194. doi:10.1080/01621459.1955.10501299 — Decomposition of a difference between two rates into composition and rate components — the method behind separating a real effect from a mix effect.
  375. 356 Mantel, N., & Haenszel, W. (1959). Statistical Aspects of the Analysis of Data From Retrospective Studies of Disease. JNCI: Journal of the National Cancer Institute, 22(4), 719–748. doi:10.1093/jnci/22.4.719 — The case-control design this page’s cohort comparisons are structured as. Crossref carries no author list, volume, issue or pages for this deposit; all four were read from the PubMed record, PMID 13655060.
  376. 357 Cohen, J. (1960). A Coefficient of Agreement for Nominal Scales. Educational and Psychological Measurement, 20(1), 37–46. doi:10.1177/001316446002000104 — Cohen’s kappa, the statistic used to test machine reading against human reading.
  377. 358 Fleiss, J. L. (1971). Measuring nominal scale agreement among many raters. Psychological Bulletin, 76(5), 378–382. doi:10.1037/h0031619 — The extension of kappa past two raters, for the case where more than one analyst judges the same span.
  378. 359 Blinder, A. S. (1973). Wage Discrimination: Reduced Form and Structural Estimates. The Journal of Human Resources, 8(4), 436. doi:10.2307/144855 — The independent derivation of the same decomposition, published in the same year. Crossref deposits a start page only.
  379. 360 Murphy, A. H. (1973). A New Vector Partition of the Probability Score. Journal of Applied Meteorology, 12(4), 595–600. doi:10.1175/1520-0450(1973)012<0595:anvpot>2.0.co;2 — The decomposition of the probability score into reliability, resolution and uncertainty — the reason a calibration number and a discrimination number are reported separately rather than folded into one figure.
  380. 361 Oaxaca, R. (1973). Male-Female Wage Differentials in Urban Labor Markets. International Economic Review, 14(3), 693. doi:10.2307/2525981 — The wage-gap decomposition that generalises Kitagawa’s method to regression, and supplies the explained/unexplained split. Crossref deposits a start page only.
  381. 362 Landis, J. R., & Koch, G. G. (1977). The Measurement of Observer Agreement for Categorical Data. Biometrics, 33(1), 159–174. doi:10.2307/2529310 — The agreement benchmarks a kappa value is read against. Crossref deposits a start page only; the range 159–174 was confirmed from the PubMed record, PMID 843571.
  382. 363 Efron, B. (1979). Bootstrap Methods: Another Look at the Jackknife. The Annals of Statistics, 7(1). doi:10.1214/aos/1176344552 — The bootstrap — the resampling procedure behind every uncertainty band on a ranked score that has no closed-form distribution. Crossref carries no page range for this deposit.
  383. 364 Rosenthal, R. (1979). The file drawer problem and tolerance for null results.. Psychological Bulletin, 86(3), 638–641. doi:10.1037/0033-2909.86.3.638 — The file-drawer problem, and the arithmetic of how many unpublished nulls would overturn a published effect — against reading any rate measured over the published record, including the rate of stated limitations, as a rate in the underlying research.
  384. 365 DeGroot, M. H., & Fienberg, S. E. (1983). The Comparison and Evaluation of Forecasters. The Statistician, 32(1/2), 12. doi:10.2307/2987588 — The formal definition of a well-calibrated forecaster, and of refinement as the thing calibration alone does not buy — what a reliability diagram is a picture of. Crossref deposits a start page only.
  385. 366 DeLong, E. R., DeLong, D. M., & Clarke-Pearson, D. L. (1988). Comparing the Areas under Two or More Correlated Receiver Operating Characteristic Curves: A Nonparametric Approach. Biometrics, 44(3), 837. doi:10.2307/2531595 — The nonparametric test for a difference between two ROC curves computed on the same cases — the correct comparison when two classifiers are scored on one corpus rather than two. Crossref deposits a start page only.
  386. 367 Church, K. W., & Hanks, P. (1990). Word Association Norms, Mutual Information, and Lexicography. Computational Linguistics, 22–29. aclanthology.org/J90-1003/ — Pointwise mutual information as a collocation statistic — the association measure behind the cue phrases the limitation detector keys on. No DOI is registered for this article; the record was read from the ACL Anthology BibTeX deposit for the Computational Linguistics version.
  387. 368 Cicchetti, D. V., & Feinstein, A. R. (1990). High agreement but low kappa: II. Resolving the paradoxes. Journal of Clinical Epidemiology, 43(6), 551–558. doi:10.1016/0895-4356(90)90159-m — The companion paper, which resolves the paradoxes by reporting kappa beside two marginal-adjusted variants — the live disagreement about what an agreement number should even be.
  388. 369 Feinstein, A. R., & Cicchetti, D. V. (1990). High agreement but low Kappa: I. the problems of two paradoxes. Journal of Clinical Epidemiology, 43(6), 543–549. doi:10.1016/0895-4356(90)90158-l — Two coders can agree on almost every item and still score a near-zero kappa when the marginals are skewed — cited against the kappa figures reported here, because a stated limitation is exactly the rare, skewed category the paradox bites hardest on.
  389. 370 Benjamini, Y., & Hochberg, Y. (1995). Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. Journal of the Royal Statistical Society Series B: Statistical Methodology, 57(1), 289–300. doi:10.1111/j.2517-6161.1995.tb02031.x — The false discovery rate correction applied in the counterfactual.
  390. 371 Agresti, A., & Coull, B. A. (1998). Approximate is Better than “Exact” for Interval Estimation of Binomial Proportions. The American Statistician, 52(2), 119–126. doi:10.1080/00031305.1998.10480550 — The adjusted-Wald alternative and the coverage comparison behind choosing between it and the score interval.
  391. 372 Newcombe, R. G. (1998). Interval estimation for the difference between independent proportions: comparison of eleven methods. Statistics in Medicine, 17(8), 873–890. doi:10.1002/(SICI)1097-0258(19980430)17:8<873::AID-SIM779>3.0.CO;2-I — The hybrid-score interval for a difference between two independent proportions — the method used for every cohort comparison on this site.
  392. 373 Buckley, C., & Voorhees, E. M. (2000). Evaluating evaluation measure stability. Proceedings of the 23rd annual international ACM SIGIR conference on Research and development in information retrieval, 33–40. doi:10.1145/345508.345543 — Measures how much of a difference between two ranked systems survives a change of topic set — the error bar that belongs on any claim that one ranking of candidates beats another.
  393. 374 Brown, L. D., Cai, T. T., & DasGupta, A. (2001). Interval Estimation for a Binomial Proportion. Statistical Science, 16(2), 101–133. doi:10.1214/ss/1009213286 — Documents the coverage failure of the normal approximation at small n — why the Wilson interval, not the Wald one. Crossref carries no page range for this deposit; 101–133 was read from the zbMATH Open record for the same DOI.
  394. 375 Agresti, A. (2002). Categorical Data Analysis. Wiley Series in Probability and Statistics. Wiley. doi:10.1002/0471249688 — The standard reference for the categorical-data methods used throughout: contingency tables, odds ratios, and the associated intervals.
  395. 376 Zadrozny, B., & Elkan, C. (2002). Transforming classifier scores into accurate multiclass probability estimates. Proceedings of the eighth ACM SIGKDD international conference on Knowledge discovery and data mining, 694–699. doi:10.1145/775047.775151 — Isotonic-regression calibration of classifier scores — the method used to turn the sentence classifier’s output into a number that can be read as a probability.
  396. 377 Kraskov, A., Stögbauer, H., & Grassberger, P. (2004). Estimating mutual information. Physical Review E, 69(6), 066138. doi:10.1103/physreve.69.066138 — The nearest-neighbour estimator for mutual information between continuous variables — how a dependence between a calibrated score and an outcome is measured without binning it first.
  397. 378 Niculescu-Mizil, A., & Caruana, R. (2005). Predicting good probabilities with supervised learning. Proceedings of the 22nd international conference on Machine learning - ICML '05, 625–632. doi:10.1145/1102351.1102430 — The empirical comparison of Platt scaling and isotonic regression across learners — the evidence for which calibrator to fit, and for how much held-out data each of them needs.
  398. 379 Gneiting, T., & Raftery, A. E. (2007). Strictly Proper Scoring Rules, Prediction, and Estimation. Journal of the American Statistical Association, 102(477), 359–378. doi:10.1198/016214506000001437 — Strictly proper scoring rules, and why a metric that rewards shading a forecast is the wrong one to optimise — the criterion any confidence score reported here has to satisfy.
  399. 380 Hayes, A. F., & Krippendorff, K. (2007). Answering the Call for a Standard Reliability Measure for Coding Data. Communication Methods and Measures, 1(1), 77–89. doi:10.1080/19312450709336664 — Krippendorff’s alpha, and the case for a reliability coefficient that tolerates missing judgements and more than two coders — the agreement statistic used where the analyst ledger is incomplete.
  400. 381 Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-Positive Psychology. Psychological Science, 22(11), 1359–1366. doi:10.1177/0956797611417632 — Undisclosed flexibility in collection and analysis makes it trivial to present a false hypothesis as significant — against treating a limitation stated in a paper as a reliable signal about what was actually done. Crossref deposits the short title only; the subtitle is not carried in the deposit. The registrant deposits the main title only; the subtitle after the colon is not in the record, so it is not printed here.
  401. 382 Franco, A., Malhotra, N., & Simonovits, G. (2014). Publication bias in the social sciences: Unlocking the file drawer. Science, 345(6203), 1502–1505. doi:10.1126/science.1255484 — Followed a full registered cohort of studies and found the null results largely never written up, let alone published — the direct measurement of the file-drawer effect against a known denominator.
  402. 383 Pakdaman Naeini, M., Cooper, G., & Hauskrecht, M. (2015). Obtaining Well Calibrated Probabilities Using Bayesian Binning. Proceedings of the AAAI Conference on Artificial Intelligence, 29(1). doi:10.1609/aaai.v29i1.9602 — Bayesian binning into quantiles, and the binned expected calibration error that came with it — the source of the estimator whose bias later work disputes. The AAAI proceedings deposit carries no page range.
  403. 384 Benjamin, D. J., Berger, J. O., Johannesson, M., et al. (2017). Redefine statistical significance. Nature Human Behaviour, 2(1), 6–10. doi:10.1038/s41562-017-0189-z — Seventy-two statisticians arguing that the conventional threshold is far weaker evidence than it is read as — the reason Orion must not treat a significant result reported in a paper as a settled constraint. Crossref records the online-first year 2017; the issue is dated 2018.
  404. 385 Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. (2017). On Calibration of Modern Neural Networks. arXiv. doi:10.48550/arXiv.1706.04599 — Modern high-accuracy networks are systematically overconfident, and temperature scaling largely fixes it — the reason a calibration stage sits between classification and ranking here.
  405. 386 Dror, R., Baumer, G., Shlomov, S., & Reichart, R. (2018). The Hitchhiker’s Guide to Testing Statistical Significance in Natural Language Processing. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1383–1392. doi:10.18653/v1/p18-1128 — A decision procedure for choosing a significance test, together with the finding that reported NLP results are routinely published without one — cited against this project, whose classifier comparisons are exactly that kind of result.
  406. 387 Kumar, A., Liang, P., & Ma, T. (2019). Verified Uncertainty Calibration. arXiv. doi:10.48550/arXiv.1909.10155 — Shows the usual binned estimate of calibration error is biased downwards, so a model can be reported as better calibrated than it is — cited against the calibration figures this pipeline publishes.
  407. 388 Vaicenavicius, J., Widmann, D., Andersson, C., Lindsten, F., Roll, J., & Schön, T. B. (2019). Evaluating model calibration in classification. arXiv. doi:10.48550/arXiv.1902.06977 — The measured calibration error depends on the binning scheme, so two defensible choices return different verdicts on the same model — against reading any single calibration number as a property of the classifier.
  408. 389 Chicco, D., & Jurman, G. (2020). The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation. BMC Genomics, 21(1), 6. doi:10.1186/s12864-019-6413-7 — Matthews correlation as the honest single number on imbalanced data, with worked cases where F1 and accuracy flatter a classifier that has learned almost nothing — against reporting a headline F1 for a task whose positive class is this rare.
  409. 390 Roelofs, R., Cain, N., Shlens, J., & Mozer, M. C. (2020). Mitigating Bias in Calibration Error Estimation. arXiv. doi:10.48550/arXiv.2012.08668 — Quantifies the bias in binned calibration error across many trained models and supplies a corrected estimator — the measurement of how far wrong the naive number is. Cited as the arXiv preprint deposit registered with DataCite.
  410. 391 Søgaard, A., Ebert, S., Bastings, J., & Filippova, K. (2021). We Need To Talk About Random Splits. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, 1823–1832. doi:10.18653/v1/2021.eacl-main.156 — Random splits inflate measured performance relative to any realistic deployment split — against reading a held-out F1 as an estimate of what the limitation detector will do on next year’s papers.
  411. 392 Pawel, S., & Held, L. (2022). The Sceptical Bayes Factor for the Assessment of Replication Success. Journal of the Royal Statistical Society Series B: Statistical Methodology, 84(3), 879–911. doi:10.1111/rssb.12491 — A formal criterion for when a replication has actually succeeded, sceptical prior and all — the statistical answer to the question a ranked constraint eventually forces: how much new evidence would settle it.
  412. Scholarly metadata, identifiers and infrastructureThe registries this page resolves against, and the identifier schemes it depends on.
  413. 393 arXiv. Understanding the arXiv identifier. arXiv help documentation. Identifiers take the form arXiv:YYMM.NNNNN; the sequence number widened from four digits to five in January 2015. info.arxiv.org/help/arxiv_identifier.html — The scheme every arXiv identifier printed in Findings and in the worked cases is checked against.
  414. 394 DOI Foundation. DOI Handbook. doi.org/the-identifier/resources/handbook/ — The normative description of the DOI system, including the handle resolution every entry in this panel was verified through.
  415. 395 Crossref. Crossref REST API. API reference and interactive documentation. api.crossref.org/swagger-ui/index.html — The interface used to read each registrant’s own deposited record, which is where every volume, issue and page range in this panel came from.
  416. 396 European Space Agency. Hipparcos catalogues. ESA Cosmos science portal. cosmos.esa.int/web/hipparcos/catalogues — The mission’s own catalogue page, cited in place of the 1997 catalogue paper, which carries no DOI and could not be verified against any registry.
  417. 397 Brase, J. (2009). DataCite - A Global Registration Agency for Research Data. 2009 Fourth International Conference on Cooperation and Promotion of Information Resources in Science and Technology, 257–261. doi:10.1109/coinfo.2009.66 — DataCite, the registration agency behind the DOIs on datasets and preprints that Crossref does not carry; the second registry the resolver behind this panel falls back to.
  418. 398 Haak, L. L., Fenner, M., Paglione, L., Pentz, E., & Ratner, H. (2012). ORCID: a system to uniquely identify researchers. Learned Publishing, 25(4), 259–264. doi:10.1087/20120404 — ORCID, the identifier the author dossiers resolve employment records against rather than trusting a name match.
  419. 399 Knoth, P., & Zdrahal, Z. (2012). CORE: Three Access Levels to Underpin Open Access. D-Lib Magazine, 18(11/12). doi:10.1045/november2012-knoth — CORE, the aggregator that harvests open-access full text out of institutional repositories; one of the routes by which a PDF reaches the estate at all.
  420. 400 Needleman, M. H. (2012). NISO Z39.96-201x, JATS: Journal Article Tag Suite. Serials Review, 38(3), 213–214. doi:10.1016/j.serrev.2012.08.006 — JATS, the article XML a publisher delivers when it offers anything better than a PDF — the format that makes the whole Extract stage unnecessary wherever it happens to be available.
  421. 401 Franceschini, F., Maisano, D., & Mastrogiacomo, L. (2014). Errors in DOI indexing by bibliometric databases. Scientometrics, 102(3), 2181–2186. doi:10.1007/s11192-014-1503-4 — DOIs indexed wrongly by the bibliometric databases themselves, at rates high enough to break automated matching — a failure in the identifier layer every entry on this page depends on.
  422. 402 Brand, A., Allen, L., Altman, M., Hlava, M., & Scott, J. (2015). Beyond authorship: attribution, contribution, collaboration, and credit. Learned Publishing, 28(2), 151–155. doi:10.1087/20150211 — The CRediT contributor-role taxonomy — the controlled vocabulary an author dossier reads instead of inferring who did what from the order of the names.
  423. 403 Mongeon, P., & Paul-Hus, A. (2015). The journal coverage of Web of Science and Scopus: a comparative analysis. Scientometrics, 106(1), 213–228. doi:10.1007/s11192-015-1765-5 — Journal coverage of the two established databases compared by discipline and by language, showing systematic gaps rather than a random sample — a standing disagreement about what the literature even consists of.
  424. 404 Guha, R. V., Brickley, D., & Macbeth, S. (2016). Schema.org. Communications of the ACM, 59(2), 44–51. doi:10.1145/2844544 — Schema.org, the structured-data vocabulary publishers embed in article landing pages; the fallback source of bibliographic fields when a deposit is thin. Crossref deposits the short title without its subtitle. The registrant deposits the main title only; the subtitle after the colon is not in the record, so it is not printed here.
  425. 405 Boudry, C., & Durand-Barthez, M. (2020). Use of author identifier services (ORCID, ResearcherID) and academic social networks (Academia.edu, ResearchGate) by the researchers of the University of Caen Normandy (France): A case study. PLOS ONE, 15(9), e0238583. doi:10.1371/journal.pone.0238583 — Author identifier services measured against actual uptake across one institution and found thinly used — the reason an ORCID match cannot be assumed and a bare name match is still doing the work.
  426. 406 Hendricks, G., Tkaczyk, D., Lin, J., & Feeney, P. (2020). Crossref: The sustainable source of community-owned scholarly metadata. Quantitative Science Studies, 1(1), 414–427. doi:10.1162/qss_a_00022 — Crossref itself — the registry every DOI on this page was checked against, described by the people who run it.
  427. 407 Liu, W. (2020). Accuracy of funding information in Scopus: a comparative case study. Scientometrics, 124(1), 803–811. doi:10.1007/s11192-020-03458-w — Funding information in a major database checked against the papers themselves and found inaccurate and incomplete — the caution against reading a funder field as fact when ranking what a programme produced.
  428. 408 Lo, K., Wang, L. L., Neumann, M., Kinney, R., & Weld, D. (2020). S2ORC: The Semantic Scholar Open Research Corpus. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 4969–4983. doi:10.18653/v1/2020.acl-main.447 — S2ORC — the reference corpus for full-text scientific mining, and the comparison point for the estate this page indexes.
  429. 409 Martín-Martín, A., Thelwall, M., Orduna-Malea, E., & Delgado López-Cózar, E. (2020). Google Scholar, Microsoft Academic, Scopus, Dimensions, Web of Science, and OpenCitations’ COCI: a multidisciplinary comparison of coverage via citations. Scientometrics, 126(1), 871–906. doi:10.1007/s11192-020-03690-4 — Six citation sources compared over the same publications, and they do not agree about what cites what — a live disagreement about coverage that any count drawn from a single index inherits silently.
  430. 410 Peroni, S., & Shotton, D. (2020). OpenCitations, an infrastructure organization for open scholarship. Quantitative Science Studies, 1(1), 428–444. doi:10.1162/qss_a_00023 — OpenCitations, the open citation index the reference links behind this page can be checked against without a subscription.
  431. 411 Wang, K., Shen, Z., Huang, C., Wu, C.-H., Dong, Y., & Kanakia, A. (2020). Microsoft Academic Graph: When experts are not enough. Quantitative Science Studies, 1(1), 396–413. doi:10.1162/qss_a_00021 — Microsoft Academic Graph, OpenAlex’s predecessor and the source of much of its structure; the reference for how such an index is built and where it errs.
  432. 412 Hsiao, T. K., & Schneider, J. (2021). Continued use of retracted papers: Temporal trends in citations and (lack of) awareness of retractions shown in citation contexts in biomedicine. Quantitative Science Studies, 2(4), 1144–1169. doi:10.1162/qss_a_00155 — Retracted papers keep being cited and the citation contexts show the citing authors did not know — the evidence that a retraction does not propagate, so a constraint read out of a withdrawn paper may never be marked void.
  433. 413 Visser, M., van Eck, N. J., & Waltman, L. (2021). Large-scale comparison of bibliographic data sources: Scopus, Web of Science, Dimensions, Crossref, and Microsoft Academic. Quantitative Science Studies, 2(1), 20–41. doi:10.1162/qss_a_00112 — Five bibliographic sources matched publication by publication, with differences large enough that the choice of source changes the answer to the question being asked of it.
  434. 414 Priem, J., Piwowar, H., & Orr, R. (2022). OpenAlex: A fully-open index of scholarly works, authors, venues, institutions, and concepts. arXiv:2205.01833. doi:10.48550/arXiv.2205.01833 — OpenAlex, the open index that supplies the works on this page carrying an openalex.org identifier instead of a DOI.
  435. 415 French, A., Hendricks, G., Lammey, R., Michaud, F., & Gould, M. (2023). Open Funder Registry to transition into Research Organization Registry (ROR). Crossref. doi:10.64000/v3429-p7810 — The Open Funder Registry folding into ROR, which is why funder and organisation identifiers on an ingested record have to be read as a moving target rather than a fixed vocabulary. Crossref registers this as a dated post with its own DOI; there is no journal venue.
  436. 416 Delgado-Quirós, L., & Ortega, J. L. (2024). Completeness degree of publication metadata in eight free-access scholarly databases. Quantitative Science Studies, 5(1), 31–49. doi:10.1162/qss_a_00286 — Metadata completeness scored across eight free databases, with abstracts, affiliations and funding missing at high rates — the direct evidence that an identifier resolving is not the same as a record being usable.
  437. 417 Culbert, J. H., Hobert, A., Jahn, N., et al. (2025). Reference coverage analysis of OpenAlex compared to Web of Science and Scopus. Scientometrics, 130(4), 2475–2492. doi:10.1007/s11192-025-05293-3 — Reference coverage in OpenAlex measured against the subscription databases and found materially incomplete — the open index this project leans on is missing references at a rate that matters.
  438. 418 Rittman, M. (2025). Retraction Watch retractions now in the Crossref API. Crossref. doi:10.13003/692016 — Retraction Watch data reaching the Crossref API, which is how retraction status can be checked for a paper in the estate at ingest time rather than by hand afterwards. Crossref registers this as a dated post with its own DOI; there is no journal venue.

Wanted, and deliberately left out. Perryman, M. A. C., et al. (1997), “The HIPPARCOS Catalogue”, Astronomy & Astrophysics 323 — no DOI is registered for this article at either Crossref or DataCite and no registrant record could be read, so it is not listed and its page range is not asserted; the mission’s own catalogue page is cited in its place, and van Leeuwen (2007) supplies the re-reduced astrometry the worked case actually used. ISO 26324, the standard underlying the DOI system, is cited only through the DOI Handbook: iso.org returned HTTP 403 and the standard’s official title could not be read from the issuing body. Two further candidates were dropped during verification because their DOIs resolved to entirely different papers than the ones intended — one was replaced with its correct record, the other abandoned. Publisher sites that returned HTTP 403 to a direct fetch — SIAM and Project Euclid — were treated as bot-walls rather than dead links, and those entries were verified through the handle system and the registrant APIs instead; they carry the publisher bot-wall routed around verification tag.

Analyst activity — the assertion ledger

Every human judgement written against a paper, newest first. The ledger is append-only: nothing here was ever edited. A judgement that was later corrected stays visible, dimmed, with a pointer to the assertion that superseded it. This ledger — not the scorecard — is the authority; the scorecard is a derived projection that a reconciler re-heals from these rows whenever the scoring loop overwrites an overlay.

loads live from /api/orion/assertions/recent

reading the assertion ledger…

    Recorded, not retained

    Everything the funnel turned away, kept auditable: screened papers judged not worth pursuing, with the reason, and selected papers whose open-access copy could not be fetched, with the last HTTP code. Each row says where to get the paper.

    50 screened out · 253 not retrieved · nothing is deleted from the record

    YearPaperVerdictReasonWhere
    1987Electron Microscopy in Molecular Biology: A Practical Approach. Edited by J. Somno constraintno constraint statement locateddoi
    1970Normalization of Gene Expression by Quantitative RT-PCR in Human Cell Line: compno constraintno constraint statement locateddoi
    1968The microbiology of cryopedogenic soils of the subarctic with particular referenno constraintno constraint statement locatedopenalex
    1988Antibody Detectionno constraintno constraint statement locateddoi
    1972The elusive tradeoff: Speed vs accuracy in visual discrimination tasksno constraintno constraint statement locateddoi
    1999Utilization OF Apple Wash Treatments And Ultraviolet Light For The Elimination Ono constraintno constraint statement locatedopenalex
    1999Effects of Pre-exercise Muscle Glycogen Status on Muscle Phosphagens, Sarcoplasmno constraintno constraint statement locatedopenalex
    1990Experience with the use of global reference fields for compilation of aeromagnetno constraintno constraint statement locateddoi
    1996Biochemical applications of FT-IR spectroscopyno constraintno constraint statement locatedopenalex
    1990Otoneurological findings in workers exposed to styrene.no constraintno constraint statement locateddoi
    2005Neck dissections: radical to conservativeno constraintno constraint statement locateddoi
    2005A report with consensus statements of the International Society of Nephrology 20no constraintno constraint statement locateddoi
    2005Selection of reference genes for gene expression studies in human neutrophils byno constraintno constraint statement locateddoi
    2004Atrial fibrillation and survival in colorectal cancerno constraintno constraint statement locateddoi
    2003A Finite Mixture Distribution Model for Data Collected from Twinsno constraintno constraint statement locateddoi
    2003Hereditary Minisatellite Mutations among the Offspring of Estonian Chernobyl Cleno constraintno constraint statement locateddoi
    2003Anthropometric measurements from a cross-sectional survey of Irish free-living eno constraintno constraint statement locateddoi
    2001Exploring the limits of vanillyl-alcohol oxidaseno constraintno constraint statement locateddoi
    2001Simple simulations of DNA condensation.no constraintno constraint statement locatedopenalex
    2000Feasibility of Sanitizing Apple Field Bins to Eliminate Postharvest Pathogensno constraintno constraint statement locatedopenalex
    2000Biological Data and Metadata Initiatives at CSIRO Marine Research, Australia, Wino constraintno constraint statement locateddoi
    1980Atmospheric solar absorption measurements in the 9 to 11 mu m region using a diono constraintno constraint statement locatedopenalex
    1975Potential Biologically Active Agents: VIII. Synthesis of 2-Methyl (and Styryl)-3no constraintno constraint statement locateddoi
    1972Quantitative Determination of Each Component in Food Color Mixtures by Dual-waveno constraintno constraint statement locateddoi
    1964Paramagnetic resonance and relaxation of Ti3+ in rubidium alum.no constraintno constraint statement locatedopenalex
    2005Evaluation of three serum antibody enzyme-linked immunosorbent assays for Mycoplno constraintno constraint statement locateddoi
    2005Theta Rhythms Coordinate Hippocampal–Prefrontal Interactions in a Spatial Memoryno constraintno constraint statement locateddoi
    2004The effect of dietary chicken egg-yolk antibodies on the clinical response in weno constraintno constraint statement locateddoi
    2004Tumor antigens for cancer immunotherapy: therapeutic potential of xenogeneic DNAno constraintno constraint statement locateddoi
    2004Neurologic Course, Endocrine Dysfunction and Triplet Repeat Size in Spinal Bulbano constraintno constraint statement locateddoi
    2003Thrombocytopaenia in canine babesiosis and its clinical usefulnessno constraintno constraint statement locateddoi
    2003The impact of diagnostic uncertainty on antibiotic prescribing for pediatric resno constraintno constraint statement locatedopenalex
    2002A strategy for peripheral nerve allografting: Immunosuppression versus chimeric no constraintno constraint statement locatedopenalex
    2001Transfer across modality in perceptual implicit memoryno constraintno constraint statement locateddoi
    2001Light therapy for seasonal affective disorder in primary careno constraintno constraint statement locateddoi
    1999Recent advances in computational thermochemistry and challenges for the future.no constraintno constraint statement locatedopenalex
    1997Field sampling and flow injection strategies for trace analysis and element specno constraintno constraint statement locatedopenalex
    1997Factors influencing germination of six wetland Cyperaceaeno constraintno constraint statement locatedopenalex
    1996NMR studies of solid nitrogen-containing dyestuffsno constraintno constraint statement locatedopenalex
    1995An investigation of anion binding by acyclic metal-centred receptorsno constraintno constraint statement locateddoi
    1995Apparatus for time-resolved measurements of acoustic birefringence in particle dno constraintno constraint statement locateddoi
    1992Treatment of algae-induced tastes and odors by chlorine, chlorine dioxide and peno constraintno constraint statement locatedopenalex
    2005Effect of dried garlic powder tablets on postprandial increase in pulse wave velno constraintno constraint statement locateddoi
    2004Quality management in reference tests for the diagnosis of classical swine feverno constraintno constraint statement locateddoi
    2004Skinless V-Notched Fillet Yields of Tilapia (Oreochromis)no constraintno constraint statement locateddoi
    2005Análisis crítico de un artículo:La terapia de reemplazo con la combinación de T3rejectedA critical appraisal of another group's trial. Secondary literature is out of scope.doi
    2005Genome-wide scanning for linkage in 56 Dutch breast cancer families selected forrejectedPhrase matched inside a cited title in a meeting-abstract supplement. The surrender belongs to a different papdoi
    1987Growth and development of rats artificially reared on rats 'milk or rats' milk/mrejectedThe cost is labour (obtaining rats' milk), which no dated capability relieves.doi
    1998A connectionist model of path integration with and without a representation of drejectedThe intractability claim is cited from another author (Gallistel 1990) and this paper itself refutes it. Secondoi
    1993Organizational factors and the perception of motion in depthrejectedA physical instrumentation limit — measuring luminance — which no computational capability relieves.doi
    selected but not retrieved — 253 papers
    YearPaperCodeWhere
    2005Fast and accurate Polar Fourier transform403doi
    2005Ventilator-induced lung injury: from the bench to the bedside200doi
    2005Bosentan in inoperable chronic thromboembolic pulmonary hypertension403doi
    2005The management and prevention of thrombotic stent occlusion: reply403doi
    2005The limitation of acute necrosis in retro-patellar cartilage after a severe blun403doi
    2005Urinary Albumin Excretion and Its Relation With C-Reactive Protein and the Metab403doi
    2005Randomized trials in the treatment of Barrett’s esophagus403doi
    2005Smokeless tobacco and coronary heart disease: a 12-year follow-up study403doi
    2005Provider Adherence to Implementation of Clinical Practice Guidelines for Neuroge403doi
    2005Complex trait mapping in isolated populations: Are specific statistical methods 200doi
    2005Abstracts of presentations at the Bert Schram Young Investigators' Mini‐Symposiu403doi
    2005Stochastic Models for Horizontal Gene Transfer200doi
    2005Effect of Raloxifene on Serum Triglycerides in Women With a History of Hypertrig403doi
    2005Efficient sorting of genomic permutations by translocation, inversion and block 403doi
    2005HotKnots: Heuristic prediction of RNA secondary structures including pseudoknots403doi
    2005Species Concepts, Species Boundaries and Species Identification: A View from the403doi
    2005A multicenter evaluation of the Pan<i>Leuco</i>gating method and the use of gene403doi
    2005Lung Cancer Immunotherapy404doi
    2005Studies on inflammatory markers in the prediction of coronary heart disease risk200openalex
    2005Detection of Cryptosporidium parvum Oocysts on Fresh Vegetables and Herbs Using 403doi
    2005Molecular markers in malignant cutaneous melanoma: Gift horse or one‐trick pony?403doi
    2005Genetic algorithms for modelling and optimisation403doi
    2005Development of Diazacon™ as an Avian Contraceptive200openalex
    2005Disturbances of grip force behaviour in focal hand dystonia: evidence for a gene403doi
    2005The Mu-Opioid Receptor Polymorphism A118G Predicts Cortisol Responses to Naloxon200doi
    2005Abstinence-Induced Changes in Self-Report Craving Correlate with Event-Related f200doi
    2005Are placebo-controlled studies required in order to prove efficacy of antidepres403doi
    2005Treatment Enhances Ultradian Rhythms of CSF Monoamine Metabolites in Patients wi200doi
    2005Biosecurity and Infectious Animal Disease200openalex
    2005Identification and Mapping of Markers Linked to the Mi Gene for Root-knot Nemato200doi
    2005MATRIX-ASSISTED LASER DESORPTION/IONIZATION TIME-OF-FLIGHT MASS SPECTROMETRY OF 200openalex
    2005Is there any Role for Sentinel Node Mapping in Colorectal Cancer Staging? Person403doi
    2005Comparison of intraocular pressure measured by Pascal dynamic contour tonometry 200doi
    2005Systemic carboplatin for retinoblastoma: change in tumour size over time403doi
    2005Dependence among sites in protein and RNA evolution.200openalex
    2005Preparing for phase II/III HIV vaccine trials in Africa403doi
    2005Association between Genital Tract Cytomegalovirus Infection and Bacterial Vagino403doi
    2004Outcome of acute ST-segment elevation myocardial infarction in diabetics treated403doi
    2004The intubating laryngeal mask airway403doi
    2004RANTES G-401A polymorphism is associated with allergen sensitization and FEV1 in403doi
    2004The additive value of tirofiban administered with the high-dose bolus in the pre403doi
    2004Six month radiological and physiological outcomes in severe acute respiratory sy403doi
    2004Flux Coupling Analysis of Genome-Scale Metabolic Network Reconstructions403doi
    2004ORFeome Cloning and Systems Biology: Standardized Mass Production of the Parts F404doi
    2004Guidelines for the hormone treatment of women in the menopausal transition and b403doi
    2004Protein structure prediction using sparse dipolar coupling data403doi
    2004Quantitative Analysis of Nucleic Acids - the Last Few Years of Progressunretrievabledoi
    2004403doi
    2004Management of Meningococcemia200doi
    2004T-Cell Stimulation by Melanoma RNA-Pulsed Dendritic Cells200doi
    2004Identification of Antigen-Specific IgG in Sera from Patients with Chronic Prosta200doi
    2004TLR4/Asp299Gly, CD14/C-260T, plasma levels of the soluble receptor CD14 and the 200doi
    2004Apolipoprotein E Modulates Clearance of Apoptotic Bodies In Vitro and In Vivo, R200doi
    2004Modeling and simulation of the human δ opioid receptor403doi
    2004Metacognition, Distributed Cognition and Visual Design403openalex
    2004Introductory editorial403doi
    2004The red ear syndrome403doi
    2004Semantic feature knowledge and picture naming in dementia of Alzheimer?s type: A403doi
    2004Effects of Exogenous Leptin on Satiety and Satiation in Patients with Lipodystro403doi
    2004Survival in patients with peripheral vascular disease after percutaneous coronar403doi
    2004Higher Rates of Viral Suppression with Nonnucleoside Reverse Transcriptase Inhib403doi
    2004Why are physicians so skeptical about positive randomized controlled clinical tr200doi
    2004The Carceral Limb of the Public Body: Jail Inmates, Prisoners, and Infectious Di403doi
    2004ED50and ED95of Intrathecal Hyperbaric Bupivacaine Coadministered with Opioids fo200doi
    2004Microarray analysis in cystic fibrosis403doi
    2004Can a Commercial Diagnostic Ultrasound Device Accelerate Thrombolysis?403doi
    2004Determination of species-specific spawning distributions of commercial finfish i401doi
    2004Association between HER-2/<b> <i>neu</i> </b> and Vascular Endothelial Growth Fa403doi
    2004Clarifying the effects of adrenergic receptor polymorphisms by measuring synapti403doi
    2004Therapeutic robocat for nursing home residents with dementia: Preliminary inquir403doi
    2004Tri-Service Surgical Meeting403doi
    2004Study and Manipulation of the Salicylic Acid-Dependent Defense Pathway in Plants200openalex
    2003A Systematic Review of the Safety and Effectiveness of Fast-track Cardiac Anesth200doi
    2003The threat of predictable and unpredictable pain: Differential effects on centra403doi
    2003Racial Disparity in the Pharmacological Management of Schizophrenia403doi
    2003Efficient and accurate likelihood for iterative image reconstruction in x-ray co403doi
    2003Meta‐analysis: evaluation of adjuvant therapy after curative liver resection for403doi
    2003Antiplatelet therapy for preventing stroke and other vascular events after carot412doi
    2003Pathophysiology of meningococcal meningitis and septicaemia403doi
    2003Comparison of polymerase chain reaction assay, bacteriologic culture, and serolo403doi
    2003Progress toward the development of a bacterial vaccine vector that induces high-403doi
    2003Hprt<sup>(CAG)146</sup> mice: Age of onset of behavioral abnormalities, time cou403doi
    2003Occupancy of Agonist Drugs at the 5-HT1A Receptor200doi
    2003Preorganization and Reorganization as Related Factors in Enzyme Catalysis: The C403doi
    2003Sexually transmitted infections in Africa: single dose treatment is now affordab403doi
    2003Effect of non-steroidal anti-inflammatory drugs on risk of Alzheimer's disease: 403doi
    2003Assessment of colour vision as a screening test for sight threatening diabetic r403doi
    2003Reply: Effect of Divalproex Combined with Olanzapine or Risperidone in Patients 200doi
    2003Management of severe malaria: implications for research403doi
    2003Screening for Fragile X Syndrome: Parent attitudes and perspectives403doi
    2003Comparative Gene Prediction in Human and Mouse403doi
    2003Normal Mixture Models for Gene Cluster Identification in Two Dimensional Microar403doi
    2003Size and Reversal Learning in the Beagle Dog as a Measure of Executive Function 403doi
    2003DNA Microarray Technology for Biodiversity Inventories of Sulfate-Reducing Proka404openalex
    2003Risk, Death and Harm: The Normative Foundations of Risk Regulation403doi
    2002Surfactant Protein D Gene Polymorphism Associated with Severe Respiratory Syncyt200doi
    2002Comparison of intraosseous or intravenous infusion for delivery of amikacin sulf403doi
    2002Systems of analysis of posterior capsule opacification403doi
    2002Sexually acquired hepatitis403doi
    2002Predictive factors for suicidal ideation in patients with unresectable lung carc403doi
    2002Health profiles and health preferences of dialysis patients403doi
    2002Distribution in allele frequencies of predisposition‐to‐atopy genotypes in Chine403doi
    2002The Use of Chemotherapy in Soft-Tissue Sarcomas403doi
    2002Early quality of life benefits of icodextrin in peritoneal dialysis403doi
    2002Combined diagonal/Fourier preconditioning methods for image reconstruction in em403doi
    2002Markers of Increased Risk of Intracerebral Hemorrhage After Intravenous Recombin403doi
    2002Combining mapping and arraying: An approach to candidate gene identification403doi
    2002Nonparametric discriminant analysis of phytoplankton species using data from ana403doi
    2002Controlling and Measuring Local Composition and Properties in Lipid Bilayer Memb200doi
    2002Impaired Innate Immunity at Birth: Deficiency of Bactericidal/Permeability-Incre200doi
    2002Molecular Epidemiology of<i>Bartonella henselae</i>Infection in Human Immunodefi403doi
    2002Repetitive Transcranial Magnetic Stimulation (rTMS) in Major Depression Relation200doi
    2002A Monte Carlo Model Reveals Independent Signaling at Central Glutamatergic Synap403doi
    2002Eye movements in iconic visual search403doi
    2002To “Sniff” or Not to “Sniff”: That Is the Question200doi
    2002Are we blind to injuries in the visually impaired? A review of the literature403doi
    2002Assuring quality and performance of sustained and controlled release parenterals200doi
    2002Assessing Allele Frequencies of Single Nucleotide Polymorphisms in DNA Pools by 526doi
    2002Postexposure prophylaxis for prevention of rabies in dogs403doi
    2002A survey of STI policies and programmes in Europe: preliminary results403doi
    2002Dispersal of the Filth Fly Parasitoid<i>Spalangia cameroni</i>(Hymenoptera: Pter403doi
    2001<i>In vivo</i> platelet activation in atherothrombotic stroke is not determined 403doi
    2001Vitamin E is ineffective for symptomatic relief of knee osteoarthritis: a six mo403doi
    2001Positron Emission Tomography Compartmental Models403doi
    2001The relation between initial symptoms and signs and the prognosis of whiplash200doi
    2001Cation binding properties of calretinin, an EF-hand calcium-binding protein.200doi
    2001Antioxidant supplementation in cancer: Potential interactions with conventional 403openalex
    2001Mapping quantitative trait loci for feed efficiency in mice : a preliminary anal200openalex
    2001Markovian domain fingerprinting: statistical segmentation of protein sequences403doi
    2001Can we eliminate trachoma?403doi
    2001A Randomized Trial of Azithromycin Versus Amoxicillin for theTreatment of <i>Chl403doi
    2001Determination of the TLR4 Genotype Using Allele-Specific PCR526doi
    2001An NMR study of 2-ethyl-1-butyllithium and of 2-ethyl-1-butyllithium/lithium 2-e200doi
    2001Frequency‐domain fluorescence microscopy with the LED as a light source403doi
    2000Antioxidant Supplementation in Atherosclerosis Prevention (ASAP) study: a random403doi
    2000High-Throughput SNP Allele-Frequency Determination in Pooled DNA Samples by Kine403doi
    2000Simple Method for Preparation of Fluor/Hapten-Labeled dUTP526doi
    2000Accurate Prediction of Protein Functional Class from Sequence in the<i>Mycobacte403doi
    2000Approaches to gene identification in neuro-psychiatric and other complex disorde403doi
    2000Author’s reply403doi
    2000Rabies preexposure vaccination among veterinarians and at-risk staff403doi
    2000Efficacy of Zinc Supplementation in Reducing the Incidence of Recurrent Vaginiti403openalex
    2000Role of neutrophils in intestinal mucosal injury403doi
    2000Regional haemodynamic effects of recombinant murine or human leptin in conscious403doi
    2000"Loss of control" in alcoholism and drug addiction: A neuroscientific interpreta200doi
    2000Reassessment of Revegetation Strategies for Kaho'olawe Island, Hawai'i200doi
    2000Randomised controlled trial using social support and financial incentives for hi403doi
    2000Mab-Zap: A Tool for Evaluating Antibody Efficacy for Use in an Immunotoxin526doi
    2000Placebo revisited403doi
    2000Qualitative Block Design Performance in Epilepsy Patients403doi
    2000Cervene (Nalmefene) in Acute Ischemic Stroke403doi
    1999IRON DEFICIENCY IN CHILDREN: DETECTION AND PREVENTION403doi
    1999The UK Prospective Diabetes Study (UKPDS): clinical and therapeutic implications403doi
    1999Nongonococcal Urethritis—A New Paradigm403doi
    1999Human Immunodeficiency Virus Load in Breast Milk, Mastitis, and Mother‐to‐Child 403doi
    1999Human Immunodeficiency Virus‐Associated Dementia: Review of Pathogenesis, Prophy403doi
    1999Cognitive and behavioural characteristics of children with Smith-Magenis syndrom202openalex
    1999Bayesian adaptive estimation of psychometric slope and threshold403doi
    1999Co-occurrence of schizophrenia and obsessive compulsive disorder : a literature unretrievableopenalex
    1999Incorporating Water Harvesting into Plasticulture Production of Muskmelon200doi
    1999Measurement uncertainty: Approaches to the evaluation of uncertainties associate403doi
    1999Water Harvesting and Supplemental Irrigation for Improved Water Use Efficiency i200doi
    1999Multipoint Oligogenic Analysis of Age-at-Onset Data with Applications to Alzheim403doi
    1999Eradication of trachoma worldwide403doi
    1999Perceptual differentiation and category effects in normal object recognition403doi
    1998Risk Factors for Prosthetic Joint Infection: Case‐Control Study403doi
    1998G protein-coupled receptors in HIV and SIV entry: New perspectives on lentivirus403doi
    1998ADVERSE SELECTION IN THE MARKET FOR CROP INSURANCE202doi
    1998The Importance of Diagnostic Cytogenetics on Outcome in AML: Analysis of 1,612 P403doi
    1998Sequential clomiphene citrate and human menopausal gonadotrophin with intrauteri403doi
    1998SEXUALLY-DIMORPHIC PATTERNS OF CORTICAL ASYMMETRY, AND THE ROLE FOR SEX STEROID 403doi
    1997Tests of vaginal microbicides in the mouse genital herpes model200doi
    1997Systematic trial of pacing to prevent atrial fibrillation (STOP-AF).403doi
    1997Endogenous estrogens and breast cancer risk: the case for prospective cohort stuunretrievabledoi
    1997Reply403doi
    1997Major depressive disorder, depressive symptoms, and bilateral hearing loss in hi403doi
    1997Firing Properties of Head Direction Cells in the Rat Anterior Thalamic Nucleus: 403doi
    1997Functional localization of the system for visuospatial attention using positron 403doi
    1997Orienteering competition injuries: injuries incurred in the Finnish Jukola and V403doi
    1997Toward a new reference method for the leukocyte five-part differential403doi
    1996Is chronic respiratory failure in neuromuscular diseases worth treating?403doi
    1996New concepts in solid-state pulse electron spin resonance500doi
    1996Synthesis of Novel Amino Acids With Incorporation Into Peptides and Synthesis of403doi
    1996Use of unreduced gametes of diploid potato (Solanum tuberosum L.) for true potat200doi
    1996Quantitative Magnetic Resonance Imaging of Human Brain Development: Ages 4–18403doi
    1996Control of African Striga Species by Natural Products From Native Plants.403doi
    1996Analytical methodology for the determination of aluminium fractions in natural f202doi
    1995Regulation of insecticidal crystal protein production in <i>Bacillus thuringiens403doi
    1995THE FARM SIZE-EFFICIENCY RELATIONSHIP IN SOUTH AFRICAN COMMERCIAL AGRICULTURE202doi
    1995Transgenic plants as vaccine production systems403doi
    1995Progress Towards a Vaccine for Lyme Disease200doi
    1995Model potential in spectral representation form. Its fidelity to the frozen-core403doi
    1994Body composition from fluid spaces and density: analysis of methods. 1961.403openalex
    1994Resistance to insulin-stimulated glucose uptake and dyslipidemia in Asian Indian403doi
    1994What Practitioners Should Understand About Bovine Lymphosarcoma200openalex
    1994Herpes simplex encephalitis: long term magnetic resonance imaging and neuropsych403doi
    1994ROLE OF PERLITE IN FLUORIDE TOXICITY OF FLORAL CROPS200doi
    1994YIELD DIFFERENCES AND DNA POLYMORPHISMS WITHIN SWEETPOTATO cv. `JEWEL' CLONES200doi
    1994BULK PLANTING SYSTEM FOR LOW-COST SEEDING OF CABBAGE200doi
    1994STORAGE LIFE AND QUALITY OF BLACKBERRY FRUIT WHEN HELD AT TEMPERATURES GREATER T200doi
    1994THE TENNESSEE SNAP BEAN INDUSTRY200doi
    1994MECHANICAL IMPULSE RESPONSE AS A MEASURE OF TOMATO FRUIT MATURITY200doi
    1994Economic change and health service reform: likely impact on teaching, practice, 403doi
    1994EXPOSURE TO CHILLING TEMPERATURES REDUCES GROWTH AND YIELD OF `SUPERSTAR' MUSKME200doi
    1994UNICONAZOLE AND BENZYLAMINOPURINE INFLUENCE FLOWERING AND GROWTH OF ACALYPHA HIS200doi
    1993Is the sleep apnoea/hypopnoea syndrome inherited?403doi
    1993Analyses of growing degree-days for agriculture in Atlantic Canada401doi
    1993Photoinduced charge separation and broken symmetry in Franck-Condon excited ππ'*202doi
    1992Assembly of polypeptide and protein backbone conformations from low energy ensem403doi
    1992Crop Rotation Minimizes Losses from Corky Root in Florida Lettuce200doi
    1992Crop-based irrigation operations in the NWFP: Progress report no.2, Kharif 92 on200openalex
    1992Molecular receptors for aromatic molecules403doi
    1991Siriraj stroke score and validation study to distinguish supratentorial intracer403doi
    1991Epidemiology of seasonal and perennial rhinitis: clinical presentation and medic403doi
    1990SUSTAINABLE AGRICULTURE: POLICY REFORM IS NOT ENOUGH202doi
    1990PCR Technology: Principles and Applications for DNA Amplification403doi
    1990Treatment of sepsis with IgG in very low birthweight infants.403doi
    1989A New Measure of Health Status for Clinical Trials in Inflammatory Bowel Disease200doi
    1989lchthyophthiriasis in fish: genetic variation in resistance to infection200doi
    1988Birth characteristics of premenopausal women with breast cancer200doi
    1988Spontaneous Ca2+ release from the sarcoplasmic reticulum limits Ca2+-dependent t403doi
    1988Pest Control in the Public Interest: Crop Protection in California403doi
    1987Results of screening a large group of intercollegiate competitive athletes for c403doi
    1987Detection of Trichomonas vaginalis antigen in women by enzyme immunoassay.403doi
    1987Identifying Threshold Technical Advances Required for Incorporating Alternative 202doi
    1986A microcomputer-based image analyzer for quantitating neurite outgrowth403doi
    1986Precipitation of Solids in Electrochemical Cells200doi
    1986Electrochemistry in Liquid Crystals: Orientational Effects in Electrochemical Pr200doi
    1985Rapid Method to Obtain Bounds on Accuracies and Prediction Error Variances in Mi200doi
    1984Monte Carlo Comparison of Sire Evaluation Models in Populations Subject to Selec200doi
    1984Redox Conduction: Its Use in Electronic Devices200doi
    1984Investigation of SOCl2 Reduction by Cyclic Voltammetry and AC Impedance Measurem200doi
    1983ELECTROCHEMICAL REACTION ENGINEERING403openalex
    1983Normal and clinical haematology of captive cranes (gruiformes)403doi
    1982Oleate Desaturation in Young Winter Wheat Root Tissue403doi
    1982Radiolabelling of Chlamydia psittaci (strain guinea pig inclusion conjunctivitis403doi
    1980Effects of sleep state and feeding on cranial blood flow of the human neonate403doi
    1979The application of spectroscopy to the study of iron-containing biological molec403doi
    1978GROUND SQUIRREL AND PRAIRIE DOG CONTROL IN MONTANA403openalex
    1978Potassium Uniport and ATP Synthesis in Halobacterium halobium403doi
    1976Automatic Detection and Classification of Infestations of Crop Insect Pests and 200openalex
    1972Part I. NMR studies of dineopentyl-magnesium exchange ; Part II. Ab initio calcu200openalex
    1971Coordination and redox properties in solution202doi
    1970Aspects of the biology and growth of three species of Ectobius (Dictioptera: Bla200openalex
    1968Boerhaave after three hundred years.403doi
    1968Steroids in Germfree and Conventional Rats403doi
    1967GAS LIQUID CHROMATOGRAPHY OF WORT AND BEER CARBOHYDRATES II. QUANTITATIVE DETERM403doi
    1967A Coulogravimetric Study of the Sintered Silver Electrode in 1 Molar Potassium H200doi
    1964Nuclear magnetic double resonance apparatus for chemical shift measurements on<s200doi
    1963Electrochemistry in Pyridine200doi
    1961Observations on the Heart-Lung Preparation in the Rat403doi
    1950A Study of the Venous Pulse in Tricuspid Valve Disease403doi
    1931Price list of fruit trees, shrubs, plants, roses, iris and gladiolus bulbs : 193403doi

    Technology

    The corpus is selected deterministically from open metadata, stored content-addressed, and screened with verifiable character offsets. Validation computation executes on the Predictive Analytics Framework — optimized R with distributed and GPU floors, measured rather than quoted.

    OpenAlex selectioncontent-addressed corpus offset-verifiable extractionbottleneck classification feasibility arithmeticlocal LLM screening OpenSearchpmem provenance PAF execution · R / OpenMPI / GPU

    Caveats

    What the numbers above do not say. Read this before quoting any of them. Every caveat here is served with the analytics payload except the one marked as a page note; none of them is written into this page.

    • caveats load with the live analytics payload

    status

    live numbers: waiting for the first read

    measuring throughput — rates appear once the page has observed the counters move over at least five minutes

    Milestones reached

    milestones load from live counts

      pipeline freshness: loading live status…