“Reward better science” sounds simple until the reward must be defined. Novel findings can open new directions, rigorous methods can reduce false claims, and replications can correct the record. A field has limited attention and career credit, so strengthening one incentive can weaken another.
This study asks:
Which mixes of novelty reward, preregistration or rigour credit, and replication credit improve the reliability of a competitive research field without collapsing its discovery rate?
It is an institutional thought experiment, not an estimate of a real discipline. Its role is to make feedback loops explicit and to expose trade-offs that a verbal argument can hide.
Laboratories as evolving strategies
Each simulated laboratory carries three continuous traits: preference for novel questions, investment in methodological rigour, and willingness to replicate. Original hypotheses may be true or false. Higher novelty has lower prior truth in this declared world; higher rigour increases power and reduces false-positive rate. Published work earns credit, and successful laboratories are more likely to be imitated, with mutation.
Replications preferentially target influential but uncertain claims. They use resources that could otherwise produce new claims, but they can update the reliability of the published record.
Policy weights lie on a simplex:
where rewards novelty, rewards rigour, and rewards replication. The novelty-only corner is the baseline.
What “reliability” and “discovery” mean
Reliability is the proportion of supported original claims in the model’s record that are true in the simulator. Discovery rate counts true original findings per unit activity. Publication and replication rates are also retained.
These are model observables, not direct measures of scientific quality. A real claim can be partly true, a replication can differ in design, and publication can change belief without a binary verdict. The simplified metrics are useful only because their definitions remain fixed across policy comparisons.
The novelty-only failure case
Across 14 seeds, the novelty-only corner yields mean reliability approximately , discovery rate , mean rigour , and mean novelty . A mean-field positive predictive value calculation gives , reasonably close to the agent-based result.
That corner is not intended as a caricature of any actual journal. It tests a feedback mechanism: if credit follows surprising positive claims while costly rigour receives no direct reward, lower-rigour high-novelty strategies can reproduce institutionally even when they degrade the record.
The highest reliability is not automatically the best policy
The replication-only corner reaches reliability about , but discovery rate falls to . It spends most activity revisiting existing claims, so the record is dependable but new true findings arrive slowly.
More interesting points lie on the nondominated frontier. A mix with approximately rigour and replication achieves reliability and discovery rate . Increasing replication to one third gives reliability and discovery . A balanced rigour–replication mix gives reliability and discovery .
The surprising model-specific result is that strong rigour credit can improve both reliability and discovery relative to novelty-only reward. Better methods raise the fraction of genuine findings enough to compensate for reduced novelty. Replication then supplies additional correction, but excessive replication eventually crowds out original work.
Evolution matters
The policy does not simply change one generation of publications. It changes which laboratory strategies receive credit, which changes the strategy distribution, which changes future evidence.
This endogenous adaptation is why a static cost–benefit table is insufficient. A rule rewarding rigorous outcomes may initially reduce publication volume, then change the population toward methods that produce more true positives. Conversely, a rule can be gamed in a richer model; the present implementation does not include strategic relabeling or metric manipulation.
Checks against one-run storytelling
The analysis uses 28 policy points and 14 independent seeds per policy. Standard deviations are shown rather than smoothing away stochastic variation. Population-size sensitivity tests whether the result is driven by a very small laboratory population. Limiting cases verify that increasing power and lowering false-positive rate move positive predictive value in the expected direction.
The mean-field calculation is especially useful. It removes evolutionary selection and claim-network history, so it cannot reproduce the whole agent model. Its agreement with the broad reliability scale indicates that the simulator’s basic truth–power–false-positive arithmetic is coherent.
What the model omits
Real scientific institutions include heterogeneous fields, collaboration, prestige networks, funding constraints, career stages, selective reporting, measurement error, theory development, data reuse, and disagreement about what counts as replication. Truth is not a simulator bit visible to an evaluator.
The policy weights also assume credits are commensurable and enforceable. In practice, a nominal reward for rigour can become paperwork, and a replication incentive can favour easy targets. Goodhart-style adaptation is absent.
Accordingly, the study does not recommend a journal scorecard. It establishes a conditional mechanism:
When laboratory strategies evolve in response to publication credit, directly rewarding rigour and some replication can move the simulated field to a better reliability–discovery frontier than novelty-only credit.
Research extensions
A research-facing next version would add:
- field heterogeneity in base rates and experimental cost;
- explicit publication bias and file drawers;
- network diffusion of influential claims;
- strategic effort allocation and imperfect auditing;
- replication designs with heterogeneous validity;
- shocks that change methods or available instrumentation;
- calibration to empirical metascience summaries with held-out targets.
The current result is valuable precisely because these are listed as missing mechanisms rather than silently assumed away.
References
- P. E. Smaldino and R. McElreath, “The natural selection of bad science,” Royal Society Open Science, 2016. doi:10.1098/rsos.160384.
- L. Tiokhin et al., “Competition for priority harms the reliability of science, but reforms can help,” Nature Human Behaviour, 2021. doi:10.1038/s41562-020-01040-1.
- M. Gordon et al., “Examining the replicability of online experiments selected by a decision market,” Nature Human Behaviour, 2024. doi:10.1038/s41562-024-01879-8.
- J. P. A. Ioannidis, “Why most published research findings are false,” PLoS Medicine, 2005. doi:10.1371/journal.pmed.0020124.