Kodo · Under the hood

Methodology

Every proprietary metric on the site, documented: what it measures, how it's computed, and — just as importantly — what it doesn't. Our bias is toward luck-adjusted skill (expected stats over raw outcomes), stabilization (small samples get shrunk toward a prior, not trusted), and receipts — we grade ourselves on the accuracy page.

Kodo Stuff+

pitch quality

A pitch-quality metric scaled like every Stuff+ (100 = league average, 10 points = one standard deviation), but built on our own data and our own model.

How it's built

Unlike FanGraphs Stuff+ and tjStuff+, which regress on run value, Kodo Stuff+ is a gradient-boosted model trained to predict the probability a pitch generates a swing-and-miss from its physical shape (velocity, movement, release geometry, arsenal-relative differentials).

We model whiff instead of run value because it's empirically the learnable target: on our pitch data, per-pitch run value is ~unlearnable from shape alone (R² ≈ 0), while whiff holds out at AUC ≈ 0.64 and correlates 0.64 with each pitcher's actual swinging-strike rate. Scored per start, so it can be trended.

What it doesn't capture: It measures how much a pitch's shape misses bats — not command/location, not sequencing, not results. A great-shape pitch thrown down the middle still gets hit.

Rest-of-season projections

the core engine

Each player's expected rest-of-season rate stats, blending this season's performance with a multi-year prior.

How it's built

Stabilization is the whole idea. A stat is trusted in proportion to how quickly it stabilizes and how much sample you have: fast-stabilizing rates (K%, whiff) count heavily even early; slow ones (BABIP, ERA) get regressed hard toward the player's career/skill prior until the sample earns trust.

Rates are anchored to skill, not outcomes: ERA is pulled toward FIP/xERA, batting toward xBA/xwOBA. Volume (xPA / xIP) is schedule-aware — computed from the team's remaining games, lineup slot, and role, not a flat games-remaining guess.

What it doesn't capture: It's a projection, not a prophecy. It can't see a trade, a role change the team hasn't signaled, or an injury that hasn't happened. Playing-time assumptions are the biggest error source.

Peak exit velocity (90th-pct EV)

power leading indicator

The 90th percentile of a hitter's batted-ball exit velocities — a stable read on how hard he can hit the ball at his best.

How it's built

Computed from every tracked batted ball (min 40) as the 90th-percentile launch speed. We use the 90th percentile rather than max EV because max is a single-swing outlier that jumps around; the 90th percentile stabilizes in ~40 batted balls and is a genuine leading indicator for ISO and power output.

What it doesn't capture: It's a power-ceiling signal, not a production stat. A hitter with elite peak EV who chases and whiffs won't turn it into results — pair it with the plate-discipline read.

xBABIP (expected BABIP)

luck signal

What a hitter's BABIP should be given his contact quality — the gap vs. his actual BABIP flags luck.

How it's built

Derived from expected batting average on the same batted balls:xBABIP = (xBA·AB − HR) / (AB − K − HR). A hitter running well below his xBABIP is getting unlucky on balls in play (a buy-low), well above is riding hits that his contact quality doesn't support (regression coming). Validated as a ~11% better predictor of future BABIP than raw BABIP.

What it doesn't capture: Batted-ball luck is only part of BABIP — speed, ballpark, and defense also move it. xBABIP is the contact-quality component, not the whole story.

Starting-pitcher grades (A–D)

daily streaming

A daily A–D grade for each scheduled start, built to answer 'is this a good streaming play tonight?'

How it's built

Blends the pitcher's skill (Stuff+, velocity trend, K/BB) with the matchup: opponent strikeout rate vs. hand, park and home/away, the Vegas total, bullpen support, and — once posted — the confirmed opposing lineup. Graded publicly: A-graded starts post a meaningfully lower realized ERA and higher K/9 than D-graded ones (see the accuracy page).

What it doesn't capture: Single-start outcomes are extremely noisy — even an A-grade start busts often. The grade is an edge over many starts, not a guarantee on any one night. QS-probability calibration is an active tuning target.

Stuff-breakout flags

buy-low pitchers

Pitchers whose ERA lags their underlying stuff — flagged as forward-looking buy-lows.

How it's built

We flag arms with elite Stuff+ and good control whose ERA is still high, on the thesis that stuff + control predicts future run prevention better than a lagging ERA does. Backtested against statistically similar non-flagged pitchers: the flagged group posts a lower forward ERA — a measured edge, published on the accuracy page.

What it doesn't capture: It's a skills-vs-results bet. Some lagging ERAs are lagging for real reasons the model doesn't see (tipping, injury, a genuinely hittable pattern).

Game-totals model

run environment

A projected total runs for each game, from our own per-team run projections.

How it's built

We project each team's runs from its lineup vs. the opposing starter and bullpen, park, and conditions, then calibrate to a single game total. It's a projections tool — we publish our number and grade it against the final, and never redistribute sportsbook odds. The graded record lives on the accuracy page and game-totals page.

What it doesn't capture: A ~2-month graded sample is directional, not proof — it's within the range where variance still matters. Weather and late lineup scratches move totals after we lock our number.

Prospect rankings & MiLB stats

minors

Rest-of-season-oriented prospect ranks blending performance, level, and scouting grades.

How it's built

Performance is discounted by level (AAA > AA > A), hitters carry a harder strikeout penalty (contact is the skill that travels), and anyone with meaningful MLB PA/IP is excluded from the prospect pool. MiLB Statcast is shown as percentiles vs. peers, not absolute values.

What it doesn't capture: The Baseball Savant minors feed can't be split cleanly by level, so MiLB Statcast is an all-levels blend that runs hot in absolute terms — read it as a percentile, not a true-talent xwOBA. Scouting grades are third-party.

Have a methodology question or spot something off? We'd rather be corrected than wrong. See the accuracy page for how these hold up against reality.