Google was selected at $10M for Spirit’s data + internal software, with Mercor at $7.5M and micro1 later offering $12.5M. I’ve been working on a framework for a question that seems increasingly relevant: if proprietary data improves an AI system, how do you translate that into what the data is worth to a specific buyer? The first part is technical. We measure the causal contribution of the proprietary data under a controlled setup: same base model, adaptation method, compute/data budget and evals, then substitute in the proprietary dataset and measure the held-out difference. We call that Capability Alpha. Spirit also lets us work backwards because we have an observed price. At an illustrative 5% Google Cloud economic scope, using ~$99.1B annualized Cloud revenue, a 35.6% operating margin proxy, 3 years and a 12% discount rate, the selected $10M price implies about 0.236% persistent economic uplift to break even. The scope assumption matters a lot: 1% → ~1.18% 2.5% → ~0.47% 5% → ~0.236% 10% → ~0.12% We also ran the same reverse calculation on other reported AI-data transactions at a standardized 5% scope: NYT / Amazon: 0.58–0.73% News Corp / Meta: ≤1.20% Reddit / Google: ~1.42% The reverse calculation is basically an adapted reverse DCF. The measurement side builds on existing data attribution/valuation work. What we’re trying to formalize is the bridge between measurable model capability and buyer-specific economic value. The $50M–$200M Spirit range in the paper is a separate forward sensitivity analysis under explicit assumptions. We have not measured Spirit’s actual Capability Alpha, so the analysis can’t establish that the $10M price was too low. Interactive paper + assumptions/calculator: https://tracerml.ai/research/pricing-capability/ Curious how people here would approach the capability → economic value step. submitted by /u/Adr-740
Originally posted by u/Adr-740 on r/ArtificialInteligence
