2026 Fantasy Football Lab is Open for Business
Keywords: decision quality; outcome bias; resulting; expected value; calibration; start/sit; waiver allocation; process metrics
Fantasy football is a repeated decision problem under uncertainty. Weekly scoring is a noisy function of player role, game environment, injury status, opponent behavior, and residual luck. Managers nevertheless often treat realized points as sufficient statistics for decision quality. That substitution—commonly described as outcome bias or resulting—produces two systematic errors: abandoning high-expected-value processes after unlucky weeks and institutionalizing low-quality processes after lucky weeks.
This article distinguishes decision quality from outcome quality, adapts the six-element Decision Quality (DQ) chain from decision analysis, and specifies a measurement stack that SCIOFF-style labs can log without collapsing into the box score. Overall decision quality is defined by the weakest link among frame, alternatives, information, values, reasoning, and commitment—not by the average of those links, and not by points scored.
We treat forecast quality (calibration, role error, Brier score) and outcome quality (wins, points for, luck residuals) as distinct analytic layers. This paper provides worked examples for start/sit, FAAB, draft, and trade decisions. We propose a lock-time protocol and post-week review sequence for reproducible manager-level research.
The “good process / bad outcome” cell is therefore not a failure of method; it is the expected cost of operating in a lottery-heavy scoring environment. A lead back with a stable 18-touch role who exits on snap three belongs in that cell. A 22% snap-share dart throw who scores 28 points belongs in the opposite cell: bad process, good outcome. The point totals may look similar, but the scientific status of the two choices is different.
Judging the choice by the ending is outcome bias (Baron & Hershey, 1988). Annie Duke’s term names the same move in probabilistic games: treating outcome quality as a perfect signal of decision quality.
Over a season, this collapses the available sample to n = 1 and turns noise into doctrine.
Decision analysis distinguishes decision quality from outcome quality (Howard & Abbas; Spetzler, Meyer, & von Winterfeldt, 2016; Decision Education Foundation). A high-quality decision does not guarantee a high-quality result, but repeated high-quality decisions increase the expected value of the resulting distribution.
The DQ construct contains six elements, represented as a chain because the decision is no stronger than its weakest link. Additional analysis on a strong link does not repair a broken one. A “one hundred percent” score on a link does not imply perfection; it marks the point at which further improvement costs more than it is worth, including the opportunity cost of delay before lock.
What problem is being solved, over what horizon, and under which constraints? “Who do I like more?” is not a frame. “Maximize expected points this week given a coin-flip matchup,” “spend FAAB to buy a November role rather than a Thursday headline,” and “construct a roster that can withstand an early ACL injury” are frames. Mixing them is a specification error. Sitting a high-volume running back in a week when the team is already projected +16 to “hunt a splash” is a frame error, even if the splash occurs.
An option that is never listed cannot be chosen. Binary menus (“start A or sit A”; “claim the WR or do nothing”) are therefore usually incomplete. A complete start/sit set includes the stud, the matchup play, the handcuff rostered for that injury tree, a WR-for-RB flex swap, a stream, and the option to hold. A complete FAAB set includes the bid, a competing claim, preserved budget, and an alternative drop. Draft capital at a turn includes BPA, positional scarcity, stack, and handcuff. If the better action was never generated, DQ cannot be high.
Evaluate information at lock, not retrospectively on Monday. Load-bearing variables include role (snaps, routes, carry share, inside-the-5 share), the last official injury designation, and game environment (implied total, spread, surface/weather). A three-catch week on nine targets is not automatically a role change. Garbage-time explosiveness on eight snaps is not a promotion. Last year’s defensive ranks are stale if this year’s usage environment is funneling work to the player in question.
Objectives must be named before the result is known. Floor versus ceiling, this week versus Weeks 15–17, FAAB now versus November, beating this opponent versus maximizing season points are different utility functions, and they can rank the same player differently. An unstated trade-off is later misread as a process failure when the draw was merely unfavorable.
The manager must be able to complete the claim: “I am choosing X over Y because of role, environment, opportunity cost, and the value declared in Section 2.4.” “Projections say so,” “he burned me,” “the field is on him,” and “stacks are smart” are incomplete unless they are tied to game total, passing environment, role, and opportunity cost. Volume is the structural variable; matchup is a modifier; narrative is residual.
A ranking that is revised three times after an early London kickoff without a new official report is not a committed decision. FAAB caps that inflate after group-chat messages are not committed allocations. The change of rule must be written in advance: out, doubtful after the last report, or a documented snap-share regime change. Everything else is noise entering through the commitment link.
Frame: prioritize floor in a projected 3-point matchup. Alternatives: lead back, pass-catching committee mate, and slot WR indoors. Information: 68% of snaps and 80% of inside-the-5 carries over three games; game indoors, total 47.5. Values: floor. Reasoning: reception plus volume outrank TD variance from the change-of-pace back. Commitment: lock unless designated out. Result: hamstring injury on snap three; committee mate scores twice. DQ remains high. The update belongs to next week’s information link, not to this week’s verdict.
Start the prior week’s touchdown scorer on a 22% snap share because “he looks explosive.” He scores again. The result does not validate the process, because role was never the load-bearing variable.
Weak alternatives: “claim the breakout name.” Strong alternatives: bid 12% on the WR, bid 7% on the back who just inherited snaps, save for the midseason starter-injury base rate, or drop the third DST instead of the rookie lottery ticket. Overpaying 19% after two tests, despite having capped value at 8%, is a commitment failure, not new information.
An early-round running back who tears an ACL in Week 3 is not automatically a process miss if the pick matched board value, positional base rates, and the Age–Missed Games style risk the lab already studies. Conversely, a late dart that becomes a league-winner does not ratify reaching without a role thesis.
“Who won the name-brand battle?” is the wrong frame. The relevant frame is remaining schedule, roster holes, and contend-now versus contend-later values. Recutting a structurally fair deal after a 4-point Thursday is resulting unless snap share itself has collapsed.
Score each link from 0–100 at lock, then report the following:
· DQ profile: diagnoses which link fails. Do not average the spider.
· Weakest-link histogram identifies the manager’s recurring leak across the season.
· Commitment breaches: Sunday records lacking a new official status.
One hundred percent means “further work is not worth the delay,” not omniscience. Forecast quality therefore tracks the accuracy of inputs and reasoning over many trials.
· Role MAE: snaps, routes, and carries; less luck-contaminated than points MAE.
· Calibration: if a start is called a 65% favorite over the bench option, the empirical hit rate should approach 0.65.
· Brier score: average squared error between the stated probability and the observed result.
For the Brier score, yi = 1 if the chosen option beat the stated alternative. The measure penalizes both misranking and overconfidence.
Hindsight “optimal lineup %” / correct-decision rate is not a DQ metric. It grades a single realized path, including the bench dart that could not have been known to be optimal at lock. It may be retained as a descriptive statistic, but it cannot serve as the lab’s headline KPI.
Outcome quality includes wins, points for, and a luck residual: actual starter points minus expected points, with the same calculation for the opponent. Over a long season, the residual should drift toward zero. A Week 1 record should not update process weights.
Measure ΔEV at lock against the next-best listed alternative. Season process EV is the sum of those gaps. Unlisted alternatives cannot be credited. Toss-up flexes with ΔEV ≈ 0.4 should not be weighed the same as large, lock-time EV decisions.
1. Review process first: identify which link was weakest at lock.
2. Review outcome second: record new facts only, such as role shift, official injury window, or coaching statement.
3. Tag the case as process miss, information miss, or variance.
4. Update Layer B running calibration; do not update process doctrine from n = 1.
More muffin: draft value versus board, adding players with real roles on waivers, and starting stable volume.
More lotto: a single Thursday start, a streamed kicker/DST, one long score deciding a two-point loss, or a first game back from injury.
DQ still applies at the lotto end of the spectrum; however, the evidentiary weight assigned to the outcome must decline.
This protocol creates the record needed for the limits discussion that follows: the lab can distinguish variance from process failure only when it records the pre-lock frame, alternatives, information, values, reasoning, and commitment before the score appears.
DQ does not remove variance. It prevents variance from becoming methodology. That is the same scientific discipline SCIOFF already asks of indexes and consistency metrics: specify the construct, refuse to let a single noisy realization rewrite the construct, and validate over a sample large enough to separate signal from noise.
ΔEV inherits whatever projection system the manager uses; garbage in remains garbage, which is why Layer B (role error and calibration) must sit beside Layer A. Synthetic or single-league logs will not generalize across scoring systems (PPR vs. standard vs. best ball). Best ball reduces weekly commitment errors and shifts the action to the draft-time frame and alternatives. Dynasty stretches the value link across years and makes the resulting after-one Thursday even less defensible.
The principal threat to this research program is Goodhart’s law. If a site leaderboard rewards hindsight optimal-lineup percentage, managers will chase last week’s dart throws. If the lab instead rewards DQmin plus a written because-clause and calibration, the incentive is to frame the week, generate options, use role at lock, name the trade-off, and stop repeatedly adjusting the lineup.
A good process with a bad week is not a failed process. It reminds us that both decisions and chance generate fantasy points. The scientific response is the same as in any noisy forecasting domain: preserve the construct, increase the sample size, calibrate the probabilities, and allow outcomes to update information—not identity.
The surest path to better season-long distributions is not a sharper narrative after Monday Night Football. It is a chain that remains intact when the ball bounces the other way.
1. Frame the week in one sentence.
2. List at least three real constructions.
3. Check role and environment as of now.
4. Name the trade-off.
5. State why the chosen option beats the next-best listed option.
6. Define what would change the decision; then stop.
7. Record DQmin. The box score is Layer C.
Baron, J., & Hershey, J. C. (1988). Outcome bias in decision evaluation. Journal of Personality and Social Psychology.
Decision Education Foundation. Principles of decision quality; six-element chain.
Duke, A. Thinking in Bets: resulting as outcome-as-signal.
Hammond, J. S., Keeney, R. L., & Raiffa, H. (1999). Smart Choices (PrOACT).
Howard, R. A., & Abbas, A. E. Decision analysis foundations.
Keeney, R. L. (1992). Value-Focused Thinking.
Spetzler, C., Meyer, J., & von Winterfeldt, D. (2016). Decision Quality: Value Creation from Better Business Decisions.
SCIOFF Lab. Site program: probability-driven analysis, consistency indexes, process over narrative (scienceoffantasyfootball.com).
Purpose: Show how a manager moves from a pre-lock decision through post-week review without confusing process quality, forecast quality, and outcome quality.
1. Decision Frame: Define the problem, horizon, and constraint before lock.
2. Generate Alternatives: List real start/sit, FAAB, draft, or trade options rather than a binary default.
3. Information Set: Record role, injury status, and game environment as known at lock.
4. Values and Trade-offs: Declare whether the decision optimizes floor, ceiling, weekly win probability, season value, or playoff leverage.
5. Reasoning: Write the because-clause connecting role, environment, opportunity cost, and stated value.
6. Commitment Rule: Specify what new information would justify changing the decision.
Measurement layer: After the pre-lock chain is recorded, score DQmin as the weakest link; track forecast quality with calibration, role error, and Brier score; then record outcome quality separately as wins, points, and luck residuals.
Post-week review loop: Review process first, outcome second. Tag the case as process miss, information miss, or variance; update calibration over a sample; and feed only validated information back into the next pre-lock decision.