The SABRhood
  • Home
  • Today
  • Leaders
  • Awards
    • MLB Award Races
    • Insane Baseball Awards
  • Intelligence
    • Story Engine
    • Players
    • Player Change Engine
    • Teams
    • Team reports
    • Matchups
    • Pitch Lab
    • Triple-A Watch
    • History
    • League Trends
  • Projections
  • Newsletter
  • Graphics
  • Research
  • Broadcast
  • About

Methodology

How The SABRhood defines, validates, and communicates its baseball metrics.

Open methods

Trust the number because you can inspect it

Every published product has a defined grain, information cutoff, sample threshold, and method label.

Canonical grain

Pitch

314,584

One row per pitch; no full PBP shipped to the browser.

Run environment

Empty, 0 outs

0.508

Empirical expected runs to the end of the inning.

Run environment

Loaded, 0 outs

2.461

Estimated from the same season-level state table.

Interpretation rule: fixed-weight wOBA is labeled as an estimate, barrel classification is labeled as a proxy, and the pressure measure is labeled as a leverage proxy. These fields are designed for transparent reporting, not to impersonate proprietary or official metrics.

Core rules

  1. Raw play-by-play remains outside the public website.
  2. The site receives compact derived products with one documented row grain.
  3. Pregame analysis never uses information from later in the game.
  4. Estimated and proxy statistics are labeled as such.
  5. Small samples receive reliability fields instead of false certainty.
  6. Automated story candidates remain subject to editorial review.

Current analytical layers

  • Canonical pitch, plate-appearance, batted-ball, and pitcher-appearance views.
  • Empirical 24-state run expectancy and linear run value.
  • Hitter and pitcher results, contact quality, and platoon splits.
  • Pitch shape, discipline, and within-plate-appearance sequencing.
  • Recent performance compared with a non-overlapping earlier baseline.
  • Six direction-aware player change z-scores paired with six full-season MLB percentiles. Positive change always means improvement from the player’s perspective; lower results allowed therefore count as better for pitchers.
  • Bullpen workload availability and exploratory pitcher-hook modeling.
  • Historical career profiles built from Lahman career totals, Hall of Fame voting, awards, All-Star selections, postseason volume, longevity, and record standing.
  • Recognition-aware anniversary, milestone, and career-record story candidates.

What “estimated wOBA” means here

The current site field is a transparent results-based wOBA calculation. It applies fixed linear weights to walks, hit-by-pitches, singles, doubles, triples, and home runs, then divides by wOBA opportunities. It is not xwOBA and does not currently use the launch-speed and launch-angle expected-contact model from the earlier research scripts.

That expected-contact work remains a separate future package layer. Until it is ported, validated, and versioned, the site will continue to say “estimated wOBA” rather than “expected wOBA” or “xwOBA.”

Plate-appearance matchup probabilities

Simulation Phase 1 assigns every confirmed batter-versus-probable-starter matchup one mutually exclusive distribution across walk, hit-by-pitch, strikeout, single, double, triple, home run, and other out.

The current shadow model follows four steps:

  1. Overall hitter rates are shrunk toward the MLB event distribution with a 200-plate-appearance prior.

  2. Overall pitcher rates are shrunk with a 250-batter-faced prior.

  3. The relevant hitter and pitcher handedness splits are each shrunk toward that player’s overall distribution with a 125-opportunity prior.

  4. Batter and pitcher distributions are combined event by event using multinomial log5:

    [ s_e = , p_{matchup,e} = ]

The final normalization is important: all possible plate-appearance outcomes sum to exactly 100%. Recent form can make only a restrained, reliability-weighted adjustment to strikeout, walk, hit, and home-run components. The model also publishes the player and split sample sizes, a matchup reliability score, its input status, and its version.

This layer is running in shadow mode. It will not replace the existing game probabilities until historical and forward-looking calibration demonstrate that it improves out-of-sample forecasts.

Plate-appearance game-state simulation

Simulation Phase 2 consumes the Phase 1 event distributions and plays each game one plate appearance at a time. The state carried between events is the inning and half-inning, number of outs, occupied bases, score, next lineup position, and workload of the defending probable starter. Each batting order continues from the hitter after the final out rather than restarting every inning.

The probable starter remains active until a sampled batters-faced target or pitch limit is reached. The replacement distribution is an availability-weighted blend of up to six active-roster relievers, excluding the scheduled starter. Each hitter-reliever probability is rebuilt from the same player, handedness, and recent-form contract used against the starter.

Runner advancement is deliberately inspectable in the first version:

  • walks and hit-by-pitches use forced advancement;
  • all runners score on home runs and triples;
  • runners on second score on 60% of singles;
  • runners on first reach third on 30% of singles and score on 55% of doubles;
  • strikeouts and other outs do not advance runners;
  • a runner begins on second in each half-inning after the ninth.

The engine directly returns game scores, winners, one-run and extra-inning rates, starter batters faced and pitch counts, bullpen plate-appearance share, and full-game hitter distributions for hits, walks, strikeouts, home runs, extra-base hits, and total bases. The initial output remains a shadow model. That initial Phase 2 output deliberately omitted double plays, sacrifice advancement, steals, individual reliever sequence, and park-sensitive baserunning. Phase 3 adds those mechanisms below while leaving defense as an explicit future input.

Phase 3 bullpen decisions and baserunning

Phase 3 replaces the aggregate bullpen handoff with a named-reliever decision. At every pitching change, each unused active reliever receives a selection score from current availability, event-weighted performance against the next three hitters, role, inning, score margin, and leverage. Scores are converted to a probability distribution and one arm is sampled. The selected pitcher remains in the game until a sampled batters-faced workload, 30-pitch ceiling, or inning-change hook is reached. The output retains both appearance probability and the reliever’s mean selection likelihood when considered.

Occupied bases now retain the responsible lineup position. This allows the simulation to credit runs, RBI, stolen bases, caught stealing, sacrifice flies, and grounded-into-double-play events to the appropriate hitter. Other outs can advance runners, score a runner from third, or become double plays through visible league-rate assumptions.

The first park layer is intentionally modest. Venues are assigned transparent expansive, quirky, compact, or neutral geometry tiers. Those tiers adjust only runner advancement and steal-attempt rates; they are not presented as empirical causal park factors. An eventual multi-season venue model can replace them only after home-team, runner-speed, batted-ball, and scoring-environment effects are separated.

Phase 3 calibration is chronological. The engine first needs 300 settled, pregame-eligible forecasts. The earliest 70% form the training block and the later 30% remain untouched for evaluation. A calibration layer enters shadow application only if it improves both win-probability Brier score and team-runs mean absolute error on that later block. Public promotion remains a separate decision.

Phase 4 empirical running game

Phase 4 reconstructs runner identity before and after every plate appearance in the private PBP store. The training table distinguishes scoring from second on a single, first-to-third on a single, scoring from first on a double, advancement on an out, double-play opportunities, and eligible windows for steals of second or third. Stolen-base events are matched to the runner and pitcher inside the same plate appearance.

Player rates are not used raw. Every runner tendency is pulled toward the league rate with an empirical-Bayes prior, while steal success receives a separate prior because attempts are much scarcer than eligible windows. Pitcher hold profiles measure how often runners try to go, then separately retain success allowed. The daily simulator uses the probable starter’s profile when it is stable and a team-level fallback otherwise.

The venue layer joins every completed game to the official MLB schedule venue. Park rates are estimated by runner-event type and shrunk toward the league before becoming bounded multipliers. These are genuine empirical inputs, but they remain descriptive: the current version has not fully separated runner mix, defense, batted-ball direction, or weather from the venue. That composition-adjusted model is the next validation step.

The granular opportunity table remains private. Public products contain only league rates, shrunk runner profiles, pitcher-hold summaries, bounded park factors, and a model card. Phase 4 stays in shadow mode and inherits the same chronological publication gate as Phase 3.

Historical recognition model

The career-significance score exists to improve editorial ordering. It is not a replacement for WAR and should never be cited as a value statistic. Hall of Fame status, major awards, All-Star selections, career volume, longevity, postseason volume, and top-ten career ranks contribute to the current version.

Lahman does not include WAR. A future WAR layer must arrive as a separately sourced, versioned table and will remain visibly labeled by source and definition.

Player Change Engine

For each qualified player, the current build compares a 14-day window with the non-overlapping portion of the season that came before it. OPS, estimated wOBA, strikeout rate, walk rate, hard-hit rate, and run value per plate appearance are each converted to a z-score across the appropriate hitter or pitcher cohort. The largest absolute z-score becomes the selected change, while the card still shows all six so the selection can be audited.

Season percentiles are calculated separately. They describe the player’s full season level among the same perspective cohort and prevent an unusual short-term change from being confused with elite overall performance. The context signal weights the size of the selected change by recent-sample reliability; it remains a reporting-priority score, not a projection.

© 2026 The SABRhood

  • Methodology

  • Glossary

  • History Desk

  • @thesabrhood