Methodology · version 0.1

Read the abstract. Mark the boundary.

The system is designed around a narrow claim: a useful signal can be produced from public metadata without inventing what the paper does not say.

01

Ingest completely

Monday editions publish only after a complete, previously unseen NBER batch has been reconciled. Identifiers are deduplicated, revisions remain visible, and failed summaries never remove a paper from search.

02

Summarize narrowly

Structured annotations cover the question, method/data, finding, importance, and caveat. Every record stores model, prompt version, generation time, source hash, and an `abstract_only` disclosure.

03

Rank blind

Author names are removed before ranking. Numerical scores remain internal. Near-duplicate questions are filtered and no more than two of the up to five selected papers may come from one primary channel.

04

Review and withhold

Low-confidence or schema-invalid annotations remain searchable as metadata but cannot enter editorial selection. Eligible machine summaries are labeled as AI-generated and unreviewed; only human-reviewed fixtures use the reviewed label. “Not stated in the public abstract” is a valid and preferred result.

Internal selection rubric

Weights guide comparison, not truth.

  1. 01Economic importance30%
  2. 02Novelty or surprise25%
  3. 03Cross-field relevance20%
  4. 04Clarity of evidence / design15%
  5. 05Topical salience10%

Public-launch gate

30–50-paper evaluation set

Required: zero unsupported numbers, methods, causal claims, or conclusions; at least 90% of summaries need no factual edit. A prompt or model change resets the evaluation.