Data Management

Health plan growth strategy: The blind spots hiding in your data

You know the read. The one at the center of your growth plan. It lays out the market you’re certain you’ll win—sized and defended with conviction. You built it carefully and grounded your recommendations in reliable, trustworthy data. But the very thing you thought was trustworthy is exactly where the biggest risk in your plan hides: in the data itself.

Key points

  • The Census stopped publishing county-to-county migration; the IRS series that replaced it counts filers, not residents, so your old model won’t transfer.
  • CMS enrollment files skip the 20 state-run exchanges, including California and New York: 7.2 million of 23 million 2026 plan selections you can’t see.
  • County business-formation data runs six to eighteen months behind the national number, so your local employer signal never lines up with the headline.
  • The one benchmark with commercial spending won’t let you download that piece, and the Dartmouth Atlas plans fall back on stopped updating in 2024.
  • Even a clean data warehouse inherits every gap in the public files feeding it, so better internal data won’t close these holes.

Is your official data complete, really?

A handful of metrics drive most growth plans: population shifts, new employer formation, per-capita spending, and enrollment by geography. Most come from public sources, and that’s fine most of the time.

But “official” isn’t the same as “complete.” Every source tells you what it includes; what matters more is what it leaves out, and that rarely gets labeled. Those missing pieces don’t announce themselves, so once the number lands in your spreadsheet, the fine print stays behind and no one downstream thinks to ask.

Let’s look at how county-to-county migration metrics have changed over the years as an example of how this happens. These metrics used to be a reliable official source for your growth planning. Before 2016, the Census Bureau’s American Community Survey published flows between individual counties, which was about as close as public data gets to watching households actually move. Then, after the 2016 to 2020 release, the Bureau stopped publishing county flows and switched to a broader state-to-county format.

The Bureau change forced many plans to turn to other official sources to obtain that information, such as the IRS. The IRS now builds county-to-county flows based on the year-to-year address changes on tax returns. It looks like a like-for-like swap for the Bureau’s data. It isn’t, and three differences show why:

  • The IRS publishing cycle takes three to four years, during which significant migration could have occurred. The IRS released its 2022 to 2023 files in March 2026.
  • The IRS is a different instrument: it counts tax filers, not residents, so it thins out exactly the low-income and older people you most need to see.
  • Your model, built for a Census series, can’t simply be pointed at the IRS one. It has to be rebuilt. Most teams never notice they have to, because nothing announces the switch.

Unfortunately, adding more datasets doesn’t fix this. A missing variable stays missing no matter how many files land next to it. Completeness isn’t something you can buy your way past; it’s a live question sitting under most of your market decisions, and it rarely shows up in the output you’re staring at.

Which gaps impact a growth decision?

Exchange sizing: a coverage gap

If you’re sizing an individual-market opportunity, CMS Marketplace enrollment at the county and ZIP level is the obvious place to start. The catch: those public-use files only cover states on the HealthCare.gov platform. The 20 states that run their own exchanges don’t show up at all.

That’s not a rounding error: in the 2026 open enrollment period, they accounted for about 7.2 million of 23 million plan selections across those states and the District of Columbia. California and New York are in that group. So is Illinois, which moved to its own platform for 2026 after a year on HealthCare.gov. In other words, a state can drop out of your file between one planning cycle and the next.

Employer formation: a timing gap

Business Formation Statistics are released monthly at the national and state levels, which makes them a useful early signal. But the county detail is annual, and it lands about six months after year-end, so your granular picture always trails the national headline by six to eighteen months, depending on where your planning falls in the cycle. That lag is manageable once you know it’s there. The trap is reading a monthly national trend and a county breakdown as if they describe the same moment: they don’t.

Spending: an access gap

The most complete multi-payer spending comparison covers Medicare, Medicaid, and commercial plans, but you can only download the first two. The private-insurance data and the combined measure can’t be posted publicly. The commercial number, the one you actually came for, is the one behind the wall. And the source itself has stopped moving. After a change in CMS research-data policy, the Dartmouth Atlas no longer calculates new annual rates; it’s now a historical archive through 2019.

Taken together, these gaps aren’t defects, and they aren’t aimed at you specifically. Each file is honest about what it covers, and the agencies that publish it made those choices for their own reasons. The data isn’t broken; it just wasn’t built for the questions you’re asking.

Why good internal data doesn’t close these gaps

A clean warehouse solves this, right? Unfortunately, no. You can have full system integration and a sharp analytics team and still inherit every gap above, unchanged and now invisible. Internal data quality and external data completeness are two different problems, and getting the first right does nothing for the second.

The provider registry is the sharpest case

The national provider registry holds roughly 6.2 million identifiers. According to a 2022 Curatus analysis, about 8.2 percent of those records were updated in the past year, and 57 percent had gone more than five years without an update. That analysis is a few years old now and comes from a provider-data vendor, so the figures are worth re-verifying before they anchor a decision. Network adequacy analysis, whether a plan has enough in-network providers close enough to members, runs on that registry. So does directory maintenance, and so does provider outreach. None of it can be repaired downstream, because the field that is wrong is wrong at the source.

The objection, and why it holds only halfway

Nobody plans on public data alone, and that’s correct. You know your own market: claims history, broker conversations, last year’s results, and a regional lead with 10 years in a territory who knows the local providers and employers better than any national file does. That closeness to the ground is a real edge, and it usually beats a spreadsheet. The catch is sequence. Those corrections arrive after the plan is funded, the network is contracted, and the targets are set. Internal evidence is good at describing where a plan has already been. The forward-looking half of the exercise still leans on the public files.

“I’ve watched plans defend a growth number with total conviction, then spend the next two years explaining what happened to it. The market doesn’t care how confident you were; it just charges you full price for finding out late.”

Jay Sivasailam, VP of Healthcare Strategy, Data Axle

Why more data isn’t the answer

Buying another dataset just adds a fourth scope to the three already in play. What changes the outcome isn’t more data, it’s connection. Most plans already have plenty of data. The real challenge is making intelligence out of it: linking signals that today sit in separate files, keeping them current, and holding them at an aggregate, de-identified level, so a migration figure and an employer-formation figure describe the same market at the same moment. Without that, coverage is just a larger pile of files.

AI raises the price of every gap above

There is a reason this problem is getting more expensive, not less. AI is only as good as the data underneath it. It doesn’t create better decisions; it amplifies the quality of the intelligence beneath it, good or bad. Point an AI model at a file missing 20 states, and it will produce a confident, fluent, wrong answer about all 50. When Deloitte’s Center for Health Solutions surveyed 100 US healthcare technology executives in September 2025, half of whom were at health plans, the organizations already getting value from AI in operations were overwhelmingly those with strong data foundations and governance in place (Deloitte, 2026).

“AI could be the best thing that ever happened to a mid-market growth plan, or the fastest way to scale a bad assumption across it. The model doesn’t decide which. The data underneath it does.”

Natalie Cunningham, SVP, Marketing, Data Axle

A strong data foundation is no longer a luxury reserved for enterprise plans with large headcounts. It’s now essential for every plan, because it decides whether AI pays you back or just scales your blind spots faster, and how much that matters depends on how much room you have to absorb a mistake. But that room isn’t the same for everyone. It scales with the size of your book, which is why the same mistake lands very differently on a national plan than on a regional one.

For a regional or mid-market plan, the real issue is exposure. A national book can absorb a bad county; a concentrated regional book cannot, and a single underpriced market can move the whole book. That’s what makes the foundation, connected, current, and complete data, the thing that decides whether an analytics or AI investment pays back, and the hardest thing to assemble file by file without a data-science bench. In an AI-driven planning cycle, these gaps don’t hold still. They compound.

Check the inputs before you defend the number

Even careful analysis inherits whatever’s missing beneath it, and public health data carries more blind spots than most plans account for. That’s how a number that looks fine gets defended in a planning meeting, only to rest on a file that stopped describing the market two releases ago. By the time the market corrects it for you, the capital’s committed and the network’s contracted.

The fix is simpler than a data-quality audit. Before a number goes into a plan, ask five things:

  1. What geography does the source actually cover?
  2. Whom does it leave out?
  3. How often is this number refreshed?
  4. Is the local number as current as the national one?
  5. When was this file last updated?

None of them are technical, and most public market data will fail at least one.

That’s what Data Axle for Healthcare is for: our platform that presents your market as a single, connected, current view at an aggregate, de-identified level, rather than a stack of mismatched files. And where a public series goes dark, fifty years of longitudinal data keeps describing what a snapshot can’t. So, the next time you defend a growth number, it’s one the market can’t quietly rewrite on you.

This is Part 1 of a three-part series on building a payer growth plan you can defend.

Part 2 is next: “The State Looks Flat. The Counties Aren’t.” It covers why sizing your market on a statewide average can cost you.

Ready to get ahead of these blind spots? See what Data Axle for Healthcare can do for you.

Brooke O’Keefe
VP, Brand & Experience

Brooke O’Keefe is a seasoned marketing strategist with a passion for blending creativity and data to drive business growth. With deep experience across both B2B and B2C brands, she brings a unique perspective to building strategies that resonate with diverse audiences and deliver measurable results. As a leader in brand, content, and go-to-market strategy, Brooke has spent her career forging strong cross-functional partnerships that connect teams, customers, and purpose. She thrives at the intersection of storytelling and technology—crafting campaigns that make brands more human and measurable.