You know the read. The one at the center of your growth plan. It lays out the market you’re certain you’ll win—sized and defended with conviction. You built it carefully and grounded your recommendations in reliable, trustworthy data. But the very thing you thought was trustworthy is exactly where the biggest risk in your plan hides: in the data itself.
A handful of metrics drive most growth plans: population shifts, new employer formation, per-capita spending, and enrollment by geography. Most come from public sources, and that’s fine most of the time.
But “official” isn’t the same as “complete.” Every source tells you what it includes; what matters more is what it leaves out, and that rarely gets labeled. Those missing pieces don’t announce themselves, so once the number lands in your spreadsheet, the fine print stays behind and no one downstream thinks to ask.
Let’s look at how county-to-county migration metrics have changed over the years as an example of how this happens. These metrics used to be a reliable official source for your growth planning. Before 2016, the Census Bureau’s American Community Survey published flows between individual counties, which was about as close as public data gets to watching households actually move. Then, after the 2016 to 2020 release, the Bureau stopped publishing county flows and switched to a broader state-to-county format.
The Bureau change forced many plans to turn to other official sources to obtain that information, such as the IRS. The IRS now builds county-to-county flows based on the year-to-year address changes on tax returns. It looks like a like-for-like swap for the Bureau’s data. It isn’t, and three differences show why:
Unfortunately, adding more datasets doesn’t fix this. A missing variable stays missing no matter how many files land next to it. Completeness isn’t something you can buy your way past; it’s a live question sitting under most of your market decisions, and it rarely shows up in the output you’re staring at.
If you’re sizing an individual-market opportunity, CMS Marketplace enrollment at the county and ZIP level is the obvious place to start. The catch: those public-use files only cover states on the HealthCare.gov platform. The 20 states that run their own exchanges don’t show up at all.
That’s not a rounding error: in the 2026 open enrollment period, they accounted for about 7.2 million of 23 million plan selections across those states and the District of Columbia. California and New York are in that group. So is Illinois, which moved to its own platform for 2026 after a year on HealthCare.gov. In other words, a state can drop out of your file between one planning cycle and the next.
Business Formation Statistics are released monthly at the national and state levels, which makes them a useful early signal. But the county detail is annual, and it lands about six months after year-end, so your granular picture always trails the national headline by six to eighteen months, depending on where your planning falls in the cycle. That lag is manageable once you know it’s there. The trap is reading a monthly national trend and a county breakdown as if they describe the same moment: they don’t.
The most complete multi-payer spending comparison covers Medicare, Medicaid, and commercial plans, but you can only download the first two. The private-insurance data and the combined measure can’t be posted publicly. The commercial number, the one you actually came for, is the one behind the wall. And the source itself has stopped moving. After a change in CMS research-data policy, the Dartmouth Atlas no longer calculates new annual rates; it’s now a historical archive through 2019.
Taken together, these gaps aren’t defects, and they aren’t aimed at you specifically. Each file is honest about what it covers, and the agencies that publish it made those choices for their own reasons. The data isn’t broken; it just wasn’t built for the questions you’re asking.
A clean warehouse solves this, right? Unfortunately, no. You can have full system integration and a sharp analytics team and still inherit every gap above, unchanged and now invisible. Internal data quality and external data completeness are two different problems, and getting the first right does nothing for the second.
The national provider registry holds roughly 6.2 million identifiers. According to a 2022 Curatus analysis, about 8.2 percent of those records were updated in the past year, and 57 percent had gone more than five years without an update. That analysis is a few years old now and comes from a provider-data vendor, so the figures are worth re-verifying before they anchor a decision. Network adequacy analysis, whether a plan has enough in-network providers close enough to members, runs on that registry. So does directory maintenance, and so does provider outreach. None of it can be repaired downstream, because the field that is wrong is wrong at the source.
Nobody plans on public data alone, and that’s correct. You know your own market: claims history, broker conversations, last year’s results, and a regional lead with 10 years in a territory who knows the local providers and employers better than any national file does. That closeness to the ground is a real edge, and it usually beats a spreadsheet. The catch is sequence. Those corrections arrive after the plan is funded, the network is contracted, and the targets are set. Internal evidence is good at describing where a plan has already been. The forward-looking half of the exercise still leans on the public files.
“I’ve watched plans defend a growth number with total conviction, then spend the next two years explaining what happened to it. The market doesn’t care how confident you were; it just charges you full price for finding out late.”
Buying another dataset just adds a fourth scope to the three already in play. What changes the outcome isn’t more data, it’s connection. Most plans already have plenty of data. The real challenge is making intelligence out of it: linking signals that today sit in separate files, keeping them current, and holding them at an aggregate, de-identified level, so a migration figure and an employer-formation figure describe the same market at the same moment. Without that, coverage is just a larger pile of files.
There is a reason this problem is getting more expensive, not less. AI is only as good as the data underneath it. It doesn’t create better decisions; it amplifies the quality of the intelligence beneath it, good or bad. Point an AI model at a file missing 20 states, and it will produce a confident, fluent, wrong answer about all 50. When Deloitte’s Center for Health Solutions surveyed 100 US healthcare technology executives in September 2025, half of whom were at health plans, the organizations already getting value from AI in operations were overwhelmingly those with strong data foundations and governance in place (Deloitte, 2026).
“AI could be the best thing that ever happened to a mid-market growth plan, or the fastest way to scale a bad assumption across it. The model doesn’t decide which. The data underneath it does.”
A strong data foundation is no longer a luxury reserved for enterprise plans with large headcounts. It’s now essential for every plan, because it decides whether AI pays you back or just scales your blind spots faster, and how much that matters depends on how much room you have to absorb a mistake. But that room isn’t the same for everyone. It scales with the size of your book, which is why the same mistake lands very differently on a national plan than on a regional one.
For a regional or mid-market plan, the real issue is exposure. A national book can absorb a bad county; a concentrated regional book cannot, and a single underpriced market can move the whole book. That’s what makes the foundation, connected, current, and complete data, the thing that decides whether an analytics or AI investment pays back, and the hardest thing to assemble file by file without a data-science bench. In an AI-driven planning cycle, these gaps don’t hold still. They compound.
Even careful analysis inherits whatever’s missing beneath it, and public health data carries more blind spots than most plans account for. That’s how a number that looks fine gets defended in a planning meeting, only to rest on a file that stopped describing the market two releases ago. By the time the market corrects it for you, the capital’s committed and the network’s contracted.
The fix is simpler than a data-quality audit. Before a number goes into a plan, ask five things:
None of them are technical, and most public market data will fail at least one.
That’s what Data Axle for Healthcare is for: our platform that presents your market as a single, connected, current view at an aggregate, de-identified level, rather than a stack of mismatched files. And where a public series goes dark, fifty years of longitudinal data keeps describing what a snapshot can’t. So, the next time you defend a growth number, it’s one the market can’t quietly rewrite on you.
This is Part 1 of a three-part series on building a payer growth plan you can defend.
Part 2 is next: “The State Looks Flat. The Counties Aren’t.” It covers why sizing your market on a statewide average can cost you.
Ready to get ahead of these blind spots? See what Data Axle for Healthcare can do for you.
Brooke O’Keefe is a seasoned marketing strategist with a passion for blending creativity and data to drive business growth. With deep experience across both B2B and B2C brands, she brings a unique perspective to building strategies that resonate with diverse audiences and deliver measurable results. As a leader in brand, content, and go-to-market strategy, Brooke has spent her career forging strong cross-functional partnerships that connect teams, customers, and purpose. She thrives at the intersection of storytelling and technology—crafting campaigns that make brands more human and measurable.