ciss-2026-ai-referee.pages.dev
Do AI Referees See What Human Referees See? Evidence from Economics Journal Submissions Olaf Stypa TU Berlin CISS 2026 22 September 2026 Webpage: ostypa.github.io Olaf Stypa (TU Berlin) CISS 2026 1
Do AI Referees See What Human Referees See? Evidence from Economics Journal Submissions Single-authored: Olaf Stypa. Motivation Economics peer review uses scarce expert time and produces heterogeneous reports. AI manuscript feedback is emerging for authors and conferences. Its substantive overlap with observed human review remains unknown in economics. Research questions RQ1:What share of recurring human concerns does one AI report recover? Additionally RQ2:Which recurring concerns does the AI report miss? RQ3:How does the measurement provider classify unmatched AI issues against the manuscript? Contribution The paper compares issue inventories derived from human referees and an AI referee report inside economics review episodes under an explicit measurement protocol. Olaf Stypa (TU Berlin) CISS 2026 2
Paper in a nutshell Question Does one AI report recover concerns raised by at least two human referees? Data and design 42 economics review episodes, 108 human reports, and one fixed AI report per episode. Compare substantive concerns within each episode. Main result 37.7% Strict recovery in the average defined bundle. GPT-5.5 measurement; six scoring runs. Interpretation AI reports recover part of the shared human signal. Most recurring concerns remain unrecovered. Scope:Content coverage in this corpus. Effects on research quality and referee time remain unmeasured. Olaf Stypa (TU Berlin) CISS 2026 3
Related literature: review capacity and AI feedback Research area Focus Scarce referee time and delays Ellison (2002); Chetty et al. (2014); Hadavand et al. (2024) Limited referee time, incentives for timely reports, and slow publica- tion in economics. AI, incentives, and review capacity Gartenberg et al. (2026) Submission growth, AI use, and incentives in a system with limited review capacity. AI review content Liang et al. (2024); Yang et al. (2026) Content overlap and differences between AI and human reviews. AI evaluation and reviewer support Pataranutaporn et al. (2025); Huang et al. (2026); Thakkar et al. (2026) Ratings of economics and finance papers; randomized AI feedback for human reviewers. This paper Issue-level comparison of AI and human reports from the same economics journal review episodes. Olaf Stypa (TU Berlin) CISS 2026 4
The available-case corpus contains 42 bundles and 108 human reports Report-count group Bundles Human reports All human- raised Consensus Zero- consensus Two human reports 19 38 329.2 35.8 3.5 Three human reports 22 66 459.3 71.3 1.5 Four human reports 1 4 32.3 5.3 0 All bundles 42 108 820.8 112.5 5 Human-consensus issues are measured issue groups raised by at least two human reports in the same manuscript bundle. Cases with zero consensus issues are reported but excluded from consensus-recall denominators. All quantities pool 6 independent measurement runs on the same inputs; counts are cross-run means and may be fractional. Scope:The sample supports a within-corpus content comparison. It does not identify population performance. Corpus Olaf Stypa (TU Berlin) CISS 2026 5
How the AI report is produced: EconReferee uses the manuscript alone Manuscript upload and run setup Text and visual extractionPDF text, tables, figures, equations Document map and classificationstructure, paper classification, agent selection Independent review agentsparallel passes over the manuscript Dependent review agentsbuild on the prior findings Verification and admissionquote checks, absence evidence, visual evidence, deduplication Consistency and synthesiscross-finding consistency, severity calibration, report synthesis AI referee report Grounding toolsmanuscript search andtyped evidence access The workflow reads the submission-stage manuscript. Independent and dependent agents examine the paper. Verification checks evidence and removes duplicate findings. Synthesis calibrates severity and creates one report. Human reports, web search, and external retrieval remain unavailable. Benchmark input:The completed report remains fixed. Olaf Stypa (TU Berlin) CISS 2026 6
How reports are compared: source-blind issue coverage Inputs Human refereereports (2 to4 per case) AI referee report Submission-stagemanuscript Measurement Issue extraction Human issueinventories Merge sameconcerns Referenceissue matrix All rows formH; rows raisedby two or morereferees formC. Issue extractionCandidate issueinventoryA Coveragegrading (full,partial, none) Recovery andoverlap metrics GroundingvalidationGroundednessmetrics Human reports form the reference issue matrix. The AI report forms the candidate issue inventory. Alignment grades full, partial, or no coverage. The manuscript supports grounding checks for unmatched AI issues only. Source labels remain hidden during measurement. GPT-5.5 repeats the complete measurement six times on identical reports. Denominators Olaf Stypa (TU Berlin) CISS 2026 7
Recurring concerns define the observable reference signal Issue group Referee A Referee B AI report Classification Exclusion restriction of the instrument ✓ ✓ ✓Recurring, recovered Attrition between survey waves✓ ✓– Recurring, missed Clarity of the mechanism section✓– – Single referee, not recovered Positioning against related literature –✓– Single referee, not recovered Construction of survey weights – –✓Addition, checked against manuscript Benchmark rule A recurring concern appears in at least two human reports. Recurrence is not truth or editorial importance. Definitions Olaf Stypa (TU Berlin) CISS 2026 8
Held-out reports provide a within-bundle human yardstick 1 Select bundles with at least three human reports. 2 Hold out one human report. 3 Identify concerns shared by two other reports. 4 Score the held-out report against those concerns. 5 Average defined scores within each bundle. Support rule Value Eligible bundles 23 Defined per run 21–22 Mean defined support 21.5 Zero denominator Undefined Bundle weight One vote Estimand Olaf Stypa (TU Berlin) CISS 2026 9
AI recovers 37.7% of recurring human concerns Primary estimand 37.7% [4.0 pp] For each defined bundle: recurring concerns recovered by AI concerns raised by at least two humans Then average bundles equally. Case weighted. Mean support: 37 defined bundles. Recovery rises with the strength of the human signal Issue weighted. Support is the mean concern count across six runs. Major concerns 47.5% (73.5) Moderate concerns 25.2% (32.8) Exactly two referees 34.9% (97.0) At least three referees 61.1% (15.5) All recurring concerns, issue weighted:38.5% [4.6 pp]over 112.5 issues. Full table Distribution Robustness Olaf Stypa (TU Berlin) CISS 2026 10
AI recovery is on the scale of one human referee 0.00 0.25 0.50 0.75 1.00 Recovery share Bundle 1 2 3 4 5 6 7 8 9 1011121314151617181920212223 Individual human refereeAI report Bars:±1 SD across runs Case-weighted comparison AI recovery:42.6% [3.6 pp] Held-out human recovery: 36.1% [4.3 pp] Each defined manuscript bundle receives one vote. Mean support: 21.5 of 23 eligible bundles. Unmatched AI issues 96.1% [0.6 pp]are provider-classified as manuscript-grounded. Issue weighted among unmatched, nonduplicate AI issues. Mean support: 491.3 of 511.2 issues. Weighting Issue categories Robustness Olaf Stypa (TU Berlin) CISS 2026 11
Recovery depends on the measurement provider and coverage rule Provider Coverage rule Recovery [SD] Defined bundles GPT-5.5 Full 37.7% [4.0 pp] 37.0 GPT-5.5 Full or partial 68.7% [4.4 pp] 37.0 Claude Sonnet 5 Full 25.3% [4.0 pp] 31.8 Case-weighted means across six scoring runs per provider. SD measures variation on fixed reports. The denominator also changes GPT-5.5 forms 112.5 recurring groups. Sonnet 5 forms 67.7 recurring groups, on average. Interpretation These are different measurements of the same reports. They do not compare AI referee systems. Recovery profile Merge audit and nested check Olaf Stypa (TU Berlin) CISS 2026 12
The evidence supports a prospective test of complementarity Measured here Content recovery in a fixed available-case corpus. Coverage against recurring human signals. Repeated scoring of identical reports. Manuscript connection for unmatched issues. Not established Population performance or ground truth. AI report-generation stability. Effects on revisions, decisions, time, or cost. Institutional complementarity or replacement readiness. Prospective test Randomize access to AI feedback before submission. Measure revisions, decisions, time, and cost. A journal partnership can connect feedback to editorial outcomes. Limits and next design Olaf Stypa (TU Berlin) CISS 2026 13
AI Referee corpus composition Human reports per bundle Bundles Human reports Two 19 38 Three 22 66 Four 1 4 Total 42 108 Available corpus Authors donated reconstructable review episodes through direct network invitations. The reports cover submissions from 2016 through 2025. Matched evidence Each bundle contains one manuscript, its human referee reports, and one fixed AI report. Back Olaf Stypa (TU Berlin) CISS 2026 1
Issue definition and strict matching rule Human reference construction 1 Extract substantive issues from each report. 2 Merge comments that describe the same problem. 3 Mark groups raised by at least two human reports. AI alignment rule 1 Extract issues from the fixed AI report. 2 Require the same manuscript target. 3 Require full coverage of the same problem. Label Operational meaning Recovered The AI report fully covers a recurring concern. Missed A recurring concern lacks full AI coverage. Addition An AI issue remains unmatched to the human issue matrix. Back Olaf Stypa (TU Berlin) CISS 2026 2
Measured human issue denominators Report group Bundles Reports All issues Recurring Zero recurring Two reports 19 38 329.2 35.8 3.5 Three reports 22 66 459.3 71.3 1.5 Four reports 1 4 32.3 5.3 0.0 All bundles 42 108 820.8 112.5 5.0 Measurement accounting Counts are means across six scoring runs. Fractional counts arise from cross-run averaging. Zero denominators Zero-recurring bundles remain in corpus descriptions. They remain outside recall denominators and do not receive zero scores. Back Olaf Stypa (TU Berlin) CISS 2026 3
Estimands, weighting, and zero denominators Case-weighted recovery bRcase = 1 |D| X i∈D ri ci Each defined manuscript bundle receives one vote. Undefined cases A bundle withc i = 0 is undefined. It does not receive a zero score. Issue-weighted recovery bRissue = P i riP i ci Bundles with more recurring concerns receive more weight. Human yardstick Hold out one report. Select common-matrix rows supported by two other reports. Average defined scores within bundles before case weighting. Back Olaf Stypa (TU Berlin) CISS 2026 4
Full AI issue-recovery table Metric Mean [SD] Defined support Case-weighted recurring concern recovery 37.7% [4.0] mean 37; range 36–38 Fixed-support recurring recovery 37.9% 35 bundles in all runs Recurring recovery, bundles with at least three reports 42.6% [3.6] 21.5 bundles Issue-weighted recurring concern recovery 38.5% [4.6] 37 bundles Case-weighted recovery of all human-raised concerns 19.9% [0.6] 42 bundles AI report overlap share 22.4% [0.8] 42 bundles Uncertainty Brackets report population standard deviations across six measurement runs on fixed inputs. They are not sampling standard errors. Back Olaf Stypa (TU Berlin) CISS 2026 5
Weighting changes the human benchmark Quantity Mean [SD] Support Weighting AI recovery on eligible bundles 42.6% [3.6] 21.5 bundles Case Held-out human recovery 36.1% [4.3] 21.5 bundles Case Descriptive AI-minus-human difference +6.6 pp [6.0] 21.5 bundles Case Fixed-support difference +4.5 pp 21 bundles in all runs Case Held-out human recovery 44.5% [4.0] 52.7 comparisons Report Case weighting Each defined manuscript bundle receives one vote. Report weighting Each defined held-out-report comparison receives one vote. Back Olaf Stypa (TU Berlin) CISS 2026 6
Case-level recovery is dispersed Recurring recovery bin Cases Share 0% to 25% 15 36.6% 25% to 50% 12 29.3% 50% to 75% 7 17.1% 75% to 100% 7 17.1% Support The table contains 41 ever-defined cases. Thirty-five cases are defined in all six runs. Six cases contribute one to three runs. Interpretation Case values are per-bundle means across the pooled runs. Zero-recurring bundles remain outside the distribution. Back Olaf Stypa (TU Berlin) CISS 2026 7
Recovery profile and measurement-provider sensitivity Recurring concern group Weight GPT-5.5 [SD] Sonnet 5 [SD] All recurring concerns Case 37.7% [4.0] 25.3% [4.0] Major Issue 47.5% [4.9] 32.9% [4.7] Moderate Issue 25.2% [7.0] 16.1% [3.9] Exactly two referees Issue 34.9% [6.1] 26.5% [2.6] At least three referees Issue 61.1% [8.3] 37.8% [9.8] Coverage-rule sensitivity for GPT-5.5 Strict full coverage gives 37.7% [4.0 pp]. The full or partial coverage rule gives 68.7% [4.4 pp]. Both estimates are case weighted over a mean of 37 bundles. Provider support and measurement boundary Mean support: GPT-5.5 / Sonnet 5 All 37.0 bundles 31.8 bundles Major 73.5 issues 47.5 issues Moderate 32.8 issues 18.5 issues Two referees 97.0 issues 61.7 issues Three or more 15.5 issues 6.0 issues Each provider rebuilds the issue inventory and denominators. GPT-5.5 forms 112.5 recurring groups. Sonnet 5 forms 67.7 recurring groups. Provider differences do not compare AI referee systems. Back to recovery Back to sensitivity Olaf Stypa (TU Berlin) CISS 2026 8
Unmatched AI issue categories Provider-classified disposition Count Unmatched share All-issue share Exact-grounded manuscript concern 92.7 18.1% 14.1% Paragraph-section grounded concern 278.8 54.6% 42.4% Thematic-grounded concern 119.8 23.4% 18.2% Generic comment 2.0 0.4% 0.3% Unsupported or off-target concern 5.3 1.0% 0.8% Unresolved anchor 12.5 2.5% 1.9% Duplicate after alignment 24.8 – 3.8% Evidence boundary Duplicates remain outside the unmatched denominator. Provider-classified grounding does not establish correctness, importance, novelty, or usefulness. Back Olaf Stypa (TU Berlin) CISS 2026 9
Merge audit and nested yardstick Hand merge audit: one run Measure Base Adj. Cons. AI recovery 33.8% 33.0% 32.3% Human yardstick 32.2% 28.1% 28.4% The audit reviewed 120 recurring groups. It scanned 33 bundles and flagged 15 missed-merge pairs. Strict nested comparison: one run Weight Human AI Diff. Case 30.3% 45.7% +15.3 pp Report 37.4% 49.2% +11.9 pp Support is 22 bundles and 52 comparisons. Each comparison rebuilds its peer-only reference. These values are not pooled estimates. Evidence boundary The audit supports the merge instrument. Back to comparison Back to sensitivity Olaf Stypa (TU Berlin) CISS 2026 10
Study limits, confidentiality, and next design Current boundary Available-case network sample. One fixed AI report per bundle. Measurement stability on fixed inputs. Provider-sensitive measurement. Unresolved threats Training-data familiarity. Possible AI help in human reports. No truth benchmark. No adoption or outcome data. Next research design Prospective journal sample. Repeated AI report generation. Blinded expert evaluation. Revision, decision, time, and cost outcomes. Confidentiality Public outputs contain aggregate evidence only. Identifiable manuscripts, reports, and issue-level audit files remain restricted. Back Olaf Stypa (TU Berlin) CISS 2026 11