Introduction
The first era of frontier AI was powered by the open web. Publicly available text, code repositories, digitized books, and captioned images pushed model capabilities far beyond what almost anyone expected. For roughly a decade, progress followed a simple formula: collect more of the data already online, train on a larger share of it, and translate scale into greater capability. The web appeared effectively free and inexhaustible, rewarding whoever could capture and process the most data.
That relationship is now beginning to shift, and the labs have said so themselves. OpenAI co-founder Ilya Sutskever told the NeurIPS conference in 2024 that “we’ve achieved peak data,” describing the internet as “the fossil fuel of AI”: a resource that formed once and has now largely been consumed. The public internet remains an important training resource and will retain relevance, but it no longer produces differentiated capability on its own. Frontier labs have trained on broadly overlapping public corpora. Future advances that separate one model from another will increasingly depend on signals the open web never contained: judgments held in experts’ minds or data sitting behind institutional walls.
The frontier is moving from abundant generic data toward scarce, capability-specific ground truth: the best available basis for defining what a model should have produced. That basis takes different forms depending on the task. In determinate clinical tasks, ground truth is operationalized through a task-specific clinical reference standard (i.e., the best available basis for defining the true condition or target for a case), which may be derived from pathology, downstream testing and follow-up, a defined clinical outcome, or adjudicated expert judgment. Open-ended generative and agentic tasks may have no single reference answer, and performance is instead assessed against clinician-authored rubrics, comparisons with qualified clinicians, or task-success criteria.
These underlying references should be distinguished from the mechanisms built on top of them. Rubrics define the criteria, human feedback supplies the judgments, graders and verifiers assess performance, and rewards provide the optimization signal. They are related components of training and evaluation systems rather than interchangeable forms of ground truth. A growing research services industry has formed to supply these inputs, organized around categories the labs themselves use: frontier training data, including expert demonstrations and preference feedback; evaluations and benchmarks; and reinforcement learning environments (i.e., task settings where a model attempts work and is scored on the result). Companies such as Scale, Surge, and Mercor built the category, and it has since expanded well beyond data labeling into post-training, evaluation, and domain-specific research services.
Frontier labs will pay increasingly large sums for ground truth. As high-quality public data becomes less differentiated, they will spend more for proprietary sources of expert judgment and verified outcomes. Supply, however, is structurally constrained: expertise takes years to develop, and outcomes can only be generated through real-world activity. This shift is occurring across coding, mathematics, science, finance, voice, and other professional domains. Healthcare is particularly important because many consequential, open-ended clinical tasks have no single, programmatically verifiable answer. Evaluating them may require qualified clinical adjudication drawing on relevant patient context, objective diagnostic evidence, and longitudinal follow-up. That is why health systems, among the most trusted institutions in healthcare, must play a critical role in shaping the next generation of health technology. We believe the defining companies of the coming decade will be those that work alongside them to bring medicine’s hard-won ground truth to the frontier. Those partnerships can also create new, diversified revenue streams for the health systems whose data and expertise make that progress possible.
The State of the Data Frontier
Scaling laws established a clean relationship between compute, data, and capability, and for most of the last decade the binding constraint was compute. Data was never trivial to assemble, and entire teams existed to gather and clean it, but it was treated as strategically abundant: the planning question was how much compute a lab could marshal, on the assumption that the web would supply whatever a larger training run required. That assumption is now under pressure from both directions.
While frontier labs are still investing heavily in scaling compute capabilities with data centers being built across the globe, the supply of public training data is growing increasingly compressed. Two forces contribute to this compression: exhaustion of high-quality public data and rising limitations on access to it.
Data exhaustion refers to a specific constraint, namely the finite supply of high-quality, public, human-generated text available for pre-training, rather than a shortage of useful data in general. That usable stock is estimated at roughly 300 trillion tokens and frontier training runs will likely exhaust it sometime before 20321. The projection carries self-disclaimed wide error bars and rests on assumptions about data efficiency that may be overturned, though the implication holds: the public stock that carried the first era of scaling is finite and is being drawn down.
Synthetic data can meaningfully extend the frontier, particularly where generated tasks can be scored using reliable feedback or formal verification (e.g., code that compiles and passes its tests). However, synthetic data’s value is limited when a capability depends on new empirical information, contextual human judgment, or outcomes that can only be observed in the real world. In those cases, a model generating its own training signal sharpens what it already represents, and it cannot supply what no one has yet recorded.

The second force is the limitation of data access. Recent research by MIT’s Data Provenance Initiative audited the consent signals on 14,000 web domains underlying the three corpora most widely used for pretraining, tracking each site’s robots.txt file (i.e., the instructions that tell automated crawlers which pages they may collect) from 2016 through April 2024. In a single year, the share of fully restricted tokens across those corpora rose from roughly 1% to as much as 7%, and among news domains restriction reached nearly 45%2. The implication for frontier labs is that restriction is rising fastest on the sources whose content is most current, most carefully curated, and most expensive to produce.

The market's response indicates how seriously the labs are competing for this signal. Surge, a data supplier that had never raised outside capital, reported revenue above $1 billion in 20243. Scale recorded roughly $870 million in revenue in 2024 and reached an annualized run rate near $1.5 billion by the end of that year4 and in June 2025 Meta invested $14.3 billion for a 49% stake in the company5 Mercor, founded in 2023, stated in June 2026 that its annualized revenue run rate had crossed $2 billion6; the company raised at a $10 billion valuation in October 20257and, as of July 2026, was reported to be in discussions at roughly $20 billion6. Mercor's figures reflect gross customer spend before payouts to the experts it contracts, so they are not directly comparable with net revenue at the other two. Even accounting for those differences, and before counting a long tail of smaller specialists, these three suppliers alone represent billions of dollars of annual spending on ground truth. The true figure is likely higher, since much post-training, evaluation, and reinforcement learning work is contracted directly between labs and specialists under private terms.
Four features of this market indicate a growing opportunity:
01. Differentiation increasingly depends on data, evaluations, and reinforcement learning environments alongside compute.
Compute remains a binding constraint, absorbing the majority of frontier capital, and is unevenly distributed across labs and geographies. Specialized data, evaluations, and reinforcement learning environments are constrained differently. Building them depends on domain expertise, institutional relationships, operating systems, and permissions that spending alone cannot immediately create. A lab can place an order for compute capacity and receive it, whereas it cannot place an order for a decade of linked imaging and outcomes at an academic medical center.
02. Useful supervision has moved up the expertise curve.
When models were weaker, generalist raters could reliably identify errors and provide corrective feedback. At the frontier, remaining errors are subtle, contextual, and domain-specific. A model that answers medical licensing questions at expert level cannot be usefully corrected by someone without clinical training, because that person cannot see what is wrong.
03. Evaluation has become a category of its own.
For most of AI research's history, measurement was an academic byproduct of the work itself. In today’s fast-moving environment, measurement has developed into a procured capability with its own vendors, methodologies, and commercial terms. Labs buy evaluation because the benchmarks that once discriminated between model generations are no longer sufficient.
04. The suppliers of this signal are increasingly specialists.
The generalist labeling model built for the computer vision era has diversified. A set of companies organized around particular domains now exists, each holding some combination of expert networks, institutional access, and measurement methodology that is difficult to assemble from scratch.
Where Ground Truth Comes From
Earlier model development relied heavily on large-scale pre-training over public and licensed data, followed by comparatively broad human feedback. As models take on more complex professional tasks, the role of humans becomes more specialized. Domain experts may be needed to author cases and adjudicate outputs in fields such as medicine or law, while representative human listeners may be better positioned to assess subjective qualities such as voice naturalness, expressiveness, and conversational fit.
Some answers can only be obtained by asking people. Whether a voice conveyed warmth, whether a response demonstrated appropriate empathy, whether a treatment plan reflects sound clinical reasoning: none of these can be tested against generic internet data. Who should be asked depends on what is being measured. Naturalness, listenability, expression, and role fit are generally best evaluated by people representative of the intended users, while clinically consequential dimensions should be defined and validated by qualified clinicians, in many cases alongside feedback from patients and caregivers. Doing either at scale, with evaluators matched to the question and scores that mean the same thing across a panel, is a discipline of its own.
Hume, an Aegis portfolio company, has built that discipline for voice. The company translates judgments about voice quality, vocal expression, and conversational behavior into structured training data and evaluation criteria. Its work spans contact centers, enterprise voice, and the frontier labs themselves, including a January 2026 partnership with Google DeepMind on voice capabilities8. Hume’s product spans three layers:
- Data: purpose-built pre-training and post-training voice corpora, annotated for emotional and vocal attributes that usage logs do not reliably capture.
- Judgment: panels screened and calibrated for the specific judgment being made, under continuous quality monitoring, built to produce reliable measurement rather than collect opinions.
- Measurement: evaluation suites built from real-world use cases, simulated conversations at scale, and head-to-head model comparisons scored against Hume's CLEAR framework (conversation, listening, expression, accuracy, reliability), grounded in roughly a decade of affective science.
Healthcare illustrates why voice-native evaluation matters. In a healthcare interaction, hesitation, pacing, or vocal strain may provide context that a transcript alone does not capture. Rather than asking whether a voice agent can independently diagnose distress, the useful evaluation asks whether it responds appropriately to what it hears and follows clinically defined escalation protocols. Some dimensions of that response can be measured automatically, though whether a response was appropriate to the vocal context still often requires human evaluation and clinical expertise.
Expert judgment is one source of ground truth. The other is the record of real-world outcomes: whether a lesion progressed, whether a patient was readmitted, whether a treatment produced the intended response. No panel can supply these answers on demand: outcome data cannot be commissioned from a rater pool at any price. It accrues over years, sits behind institutional walls, and is governed by consent, privacy law, and data use agreements that took decades to construct. Access is a function of trust, and defensibility rests on institutional relationships, longitudinal linkage across fragmented records, and traceability clean enough to withstand regulatory and scientific scrutiny.
Healthcare's Ground Truth Challenge
Healthcare is where these answers are among the most difficult to produce and where errors carry the greatest consequence. Much of medicine's most valuable signal is either never captured, fragmented across institutions and points in time, or captured without the linkage, provenance, permissions, and curation required for model development and evaluation.
The evaluation gap this produces is now measurable. A 2025 systematic review in the Journal of Medical Internet Research examined 39 AI medical benchmarks. Across those benchmarks, leading models achieved 84% to 90% on knowledge-based assessments, largely multiple-choice examinations in the style of medical licensing tests, while selected practice-based benchmarks reported success rates of 45% to 69% on tasks such as open-ended diagnostic cases, electronic health record navigation, and multiturn clinical conversations9 . Although those figures are not directly comparable, they point to a broader gap between recalling medical knowledge and performing real world clinical tasks.

The gap exists because exam questions test something models are already good at: retrieving facts that are well documented, textbook, and often available online. A question with a clean answer key rewards recall whereas real clinical cases rarely present information so cleanly. They require weighing ambiguous symptoms, accounting for a patient’s history, and recognizing when a textbook presentation is hiding something else. Such judgment depends on specialized signal that is not consistently captured in a form models can retrieve. Closing the capability gap requires giving models access to harder, more specialized signal than public sources contain.
Avandra, another Aegis portfolio company, is assembling that signal in medical imaging. A radiologist who flags a lung nodule on a chest CT is making a judgment call. Whether the call was right may only become clear through a later scan, biopsy, or other clinical follow-up, often at a different institution and recorded in a system the first hospital never sees. Avandra's federated network links those pieces together, with de-identification, governance, and provenance built into the infrastructure. Because the data never leaves institutional control, the design addresses the objection that has blocked many prior attempts. It also changes the economics of participation: health systems in the network share in revenue generated from the use of the imaging data they steward, creating a new source of revenue and helping diversify their revenue base.
Scale across institutions matters as much as access to any one of them. Clinical data carries the fingerprint of the place that produced it: scanner fleets, imaging protocols, documentation practices, and patient populations all differ, and a model validated at one center may not generalize across other scanners, protocols, populations, or care settings. Models that generalize require diversity across many sites, and a multi-institutional network can capture broader variation than most single-system datasets.
Avandra links longitudinal imaging to follow-up studies, pathology, treatment history, and outcomes, while preserving original radiologist interpretations and surrounding clinical notes as foundational labels. Depending on the task, this fuller record can support research-ready cohorts and a stronger clinical reference standard for evaluating an earlier read against subsequent clinical evidence.
Looking Ahead
Across Hume and Avandra, the scarce asset is not data volume alone, but the human judgment and longitudinal clinical context required to build models that can be trusted where trust is hardest to earn. A voice agent that hears what a patient is actually saying, and an imaging model whose calls hold up against later follow-up, become possible only when their training and their testing run through real expertise and real outcomes. The work of making that signal available is quiet compared to the model releases it enables, though it is what will separate AI that demos well from AI that practices well.
At Aegis, we believe health systems will be central architects of this next era of applied intelligence, shaping AI around the realities of care and the needs of the patients and caregivers it is meant to serve.



