Data providers: the categories, and how to pick the right one

Every data vendor sells company records. What separates them is which question those records answer, and buyers who skip that distinction end up with two contracts covering the same ground and a gap where the answer was.

This page is a map rather than a ranking. Seven categories, what each solves, and which of our detailed rankings covers it.

If you already know the category you need, skip to it. If you do not, the routing table below is the fastest way to find out.

Seven categories of data provider mapped by the subject each one describes: a company, a person, a behaviour or a technology
Seven categories, four kinds of subject. Overlap is where duplicate spend happens.

Which data vendor answers which question

Your questionCategoryTypical buyerWhere to read more
Who do I call at this company?B2B contact dataSales teams, SDRsB2B data providers
Which companies exist that match my ICP?Firmographic dataMarketing, RevOpsFirmographic data providers
What does this company run?Technographic dataSales, product, partnershipsTechnographic data providers
Who is looking to buy right now?Intent dataRevenue teams, ABMIntent data providers
What do I know about this business itself?Company dataData teams, strategy, riskCompany data providers
What can I infer before the numbers are published?Alternative dataInvestors, corporate strategyAlternative data providers
Is this entity real and who owns it?Registry and compliance dataRisk, compliance, procurementCovered inside company data providers
Important

Most teams need two or three of these rather than one. The mistake is buying two that answer the same question because both vendors described themselves as a b2b data provider.

The categories of data provider

Seven categories, in the order most teams encounter them.

B2B data providers for contact data

B2B data providers sell contact data with company records attached: names, job titles, verified email addresses, direct dials and phone verified mobile numbers, priced per seat and built for a rep to read.

This is the largest and most crowded category, and the one with the worst decay problem. B2B contact data decays at roughly 2.1% per month, 28% of email addresses go stale annually, and poor data quality costs organisations an estimated $12.9 million a year.

ZoomInfo, Cognism, Apollo, Lusha and UpLead lead it. Accuracy, regional coverage and compliance separate them far more than database size does.

Company data providers

Company data providers sell records about the business rather than the people in it: firmographics, ownership, financials, funding and technology signals, priced per volume and built for a system to consume.

The buyers are data and strategy teams rather than sales. Coresignal, Bright Data, People Data Labs, Bureau van Dijk and OpenCorporates sit here, delivering by API or bulk file rather than through a web app.

Data accuracy can deteriorate by 15% monthly across this category, which is why refresh cadence matters more than record count.

Firmographic data providers

Firmographic data is the attribute layer: industry, company size, annual revenue, geographic location, ownership type and growth indicators.

It is less a separate market than a field set every provider claims. What differs is completeness. Companies with a strong ideal customer profile see 68% higher win rates, and that ICP is built from firmographic fields that are either populated or not.

Technographic data providers

Technographic data describes technology usage: the CRM, cloud services, marketing automation and security tools a company runs.

Coverage claims vary wildly because they measure different things. ZoomInfo publishes 30,000+ technologies across 100 million companies; BuiltWith detects 123,000+ across 670 million websites; HG Insights tracks around 32,000 IT contracts with spend attached. Those are three different products wearing one label.

Intent data providers

Intent data providers sell behavioural evidence that a company is actively researching solutions in your category, gathered from publisher networks, review sites, search behaviour and your own website.

Scale claims here are enormous and largely unverifiable. Bombora tracks 17 billion interactions monthly across 5,000+ sites; Demandbase reports 500+ billion signals across 300,000 keywords; 6sense processes over a trillion buying signals daily. Treat intent signals as a prioritisation input rather than a qualification.

Alternative data providers

Alternative data means non-traditional sources read for what they reveal before disclosure: transactions, satellite imagery, web traffic, job postings, geolocation.

The buyers are investors first and operators increasingly. The global alternative data market reached $18.74 billion in 2025 with projections of $135.8 billion by 2030, and hedge funds allocate over $1.6 million a year each to it.

Registry and compliance data

The smallest category and the only one with an authoritative source. Company registries publish incorporation, officers, status and filings, and providers like OpenCorporates resell it with provenance attached to each field.

Nothing else in this list can tell you an entity legally exists. Risk, compliance and procurement teams buy here and rarely need anything else.

How to evaluate data quality

The same four checks apply whichever category you are buying, and none of them is the record count on the homepage.

Data accuracy and freshness

Ask for the accuracy rate and the verification method behind it in writing. Providers that publish one are rare enough that it is a differentiator; UpLead's 95% guarantee with credit refunds on bounces is the clearest example in this market.

Real-time verification at the point of export beats batch cleaning outright. At 2.1% monthly decay on contact data, a record verified in January is meaningfully worse by April, and continuously updated data is worth paying for on the volatile fields.

Refresh cadence varies enormously. Coresignal refreshes every six hours, Global Database updates many records within 24 hours, MixRank updates monthly while holding five years of history. Those suit an application, an operational workflow and a trend analysis respectively.

Data coverage against your ICP

Data coverage is regional and role-specific rather than a single number. A provider strong on US enterprise is often half as complete on European mid-market, and broad global coverage in the marketing copy usually means a home region plus licensed feeds.

Sub-50-employee coverage is where databases thin out. If you sell to small businesses, test there specifically, because aggregate completeness figures are carried by enterprise records.

Some vendors specialise in specific industries and beat the generalists inside them. That trade is often worth taking if your target market sits in one vertical.

Ask for a data sample

Quick tip

Never evaluate on the sample the vendor picks. Give them 50 to 100 accounts from your closed-won list, which are your ICP by definition and independently verifiable.

Score three things per record: was the company found, how many fields were populated, and how many populated fields were correct. Data depth matters as much as match rate, and the third number is the one nobody quotes.

Providers that decline to run that test are pricing on hope. The ones that run it and report their own gaps honestly are worth negotiating with.

Data sourcing and provenance

Ask which fields are observed and which are inferred. Most datasets blend both, and few vendors label the difference, which matters because you cannot weight a scoring model when every field looks equally confident.

Third party data has a structural weakness worth knowing: vendors license from each other. Two providers appearing to corroborate a figure while both sourcing it upstream from the same feed is not corroboration.

Data collected from public web sources, company websites and job boards is cheap and observable. Human verified data is accurate and does not scale. High quality data usually means knowing which you are holding.

How you access data

Delivery method decides who can use the data, and it is worth settling before price.

Web application. A person searches and exports. Suits sales teams, priced per seat, and the export is where the data goes stale.

CRM integration. Fields write onto existing records automatically. This is where data enrichment becomes infrastructure rather than a task, and integration capabilities with existing systems are what decide if the purchase survives year one.

API access. Code queries on demand. Suits applications and enrichment inside data pipelines, priced by call volume rather than seats.

Bulk datasets. Files land in a warehouse. Suits analysis, market mapping and model training, and it is the only sensible route at real data volume.

Worth checking

Paying per seat for data your application consumes is the most common overspend in this market. If code reads the record, buy the dataset.

Data compliance and security

Company records are less exposed than personal ones, and the exceptions are where teams get caught.

Contact data falls squarely under GDPR and CCPA. So do officer records, employee profiles and any dataset naming individuals, whatever the product page calls it. Data privacy regulations follow the data subject rather than the label on the file.

Check that a data provider adheres to DNC screening before phone numbers change hands, and ask how consent was obtained for European records specifically. Cognism's positioning is built on exactly this, which tells you how much buyers care.

Secure data handling is the other half. Ask where records are stored, who at the vendor can access them, and what happens to data you already pulled if the contract ends. Perpetual rights on delivered records cost more upfront and far less than re-licensing later.

Choosing the right data vendor

Four key factors, in the order that matters.

The question you need answered. Contact, company, firmographic, technographic, intent or alternative. Getting this wrong makes every other criterion irrelevant.

Coverage on your ICP. Not the market, yours. Tested on accounts you can verify.

Delivery into the tools people use. Data that requires a manual export gets applied once.

Compliance in the regions you operate. The only criterion with legal rather than commercial consequences.

Price comes fifth and still matters. Published pricing lets you model consumption; custom pricing means negotiating with less information than the vendor has.

Running more than one data provider

Most teams past their first year run several, and the pattern is consistent.

One primary provider for the category that carries the most weight. One fallback for the records the first misses, queried only on gaps, which is the waterfall enrichment approach high-performing outbound teams use.

Then a verification step before anything reaches a sequence, because accurate data at purchase is not accurate data at send.

Beyond that, one provider per distinct question. Intent alongside contact data. Technographic alongside firmographic. Adding a second vendor in the same category is duplication; adding one in a different category is coverage.

Data providers FAQ

What is a data provider?

A company that collects, structures, verifies and sells information about businesses or the people who work at them. Delivery runs from a web app through CRM integration and API access to bulk datasets, and buyers range from sales teams to model builders.

What is the difference between a data vendor and a data supplier?

In practice the terms are used interchangeably. Where a distinction is drawn, a data supplier originates the data it sells and a data vendor may resell or aggregate from others. The useful question is not the label but which fields a provider collects first-hand.

Who are the biggest data providers?

By coverage claims: ZoomInfo publishes over 500 million contacts, Coresignal 75 million company and 839 million employee profiles, BuiltWith detection across 670 million websites. Size and usefulness are separate questions, and the largest provider is rarely the right one for a specific ICP.

How much do data providers cost?

Published figures span $39 a month for narrow tools to $250,000 a year at the top of the enterprise ABM platforms. Contact data platforms commonly land between $15,000 and $40,000 annually; company data and alternative data price by volume rather than seats.

How do data providers collect their data?

Public web sources including company websites and job boards, licensed third-party datasets, contributor networks where users share their own records, registry and government filings, and direct human verification for premium fields. Most large providers blend all five.

Is buying B2B data legal?

Yes, within the rules of the jurisdiction you prospect into. GDPR governs European personal data, CCPA covers California, and phone outreach requires DNC screening. Legality depends on the provider's data sourcing and on your own usage, and both need checking.

What is the difference between a data provider and a sales intelligence platform?

A data provider sells records. A sales intelligence platform wraps records in workflow: sequencing, scoring, alerting and reporting. ZoomInfo, Apollo and Cognism sit in both categories, which is why they appear on several of our rankings under different orders.

Bottom line

Start from the question rather than the vendor list. Contact data to reach someone, firmographic to find the right companies, technographic to know what they run, intent to know when, company data to build something, alternative data to see before disclosure.

Then test on accounts you already know, check the delivery method against who will consume the records, and settle compliance before price.

Every provider in every one of these categories looks complete in a demo. The difference shows up on your own list.

What each data type contains

Six data types run through these categories, and the same record can carry several.

Contact data points. Names, job titles, verified email addresses, direct dials and mobile numbers. The most perishable type and the one every outbound motion depends on.

Firmographic data points. Industry, company size, revenue, headquarters, ownership type. The layer segmentation runs on.

Technographic. Technology usage across CRM, cloud, analytics and security tools, mostly inferred rather than observed.

Behavioral data. Content consumption, search activity and site visits. Engagement data of this kind is what intent products are built from, and it ages fastest of everything here.

Event data. Funding events, leadership changes, acquisitions and job postings. Discrete, dated and far more actionable than a static attribute.

Financial and registry data. Revenue, credit indicators, ownership chains and incorporation records, thin on private companies and authoritative where filings exist.

Internal data against external data

Everything above is external data, bought to describe companies you do not control. Internal data is what your own systems already hold: closed-won accounts, product usage, support tickets, sales notes.

The two are worth more joined than separately. Internal data defines the ICP; external data finds more companies matching it. Teams that buy external data without first mining what they already have usually buy the wrong segment.

Historical data matters for both. A single snapshot tells you what is; a time series tells you what changed, and change is what triggers outreach.

Web data as a source layer

Web data underpins most of these categories rather than sitting beside them. Company websites, job boards, review sites and public filings are where the majority of commercially sold B2B records originate.

That has a consequence worth internalising. If a fact is not published somewhere public, somebody either surveyed for it, bought it from a contributor, or estimated it. Private company data is mostly the third.

How teams use data providers

The same records serve several jobs, and the buying case usually rests on one while three others quietly benefit.

Sales and marketing

Sales and marketing teams are the largest buyer group. Contact data reaches key decision makers, firmographics build the target accounts list, intent orders it, and technographics shape the message.

Sales and marketing strategies break when the two functions read different fields. Marketing segmenting on an attribute sales cannot see produces leads reps cannot contextualise, which is a data problem presented as an alignment problem.

Sales engagement tools consume the output. Data that cannot reach a sequence without a manual export reaches it late or not at all.

Market research and market mapping

Market research uses the same firmographic records to count rather than to contact. Market mapping asks how many companies exist at what size in which geography, which is a query once the data is in place and a guess before it.

Market trends emerge from the time series. Counting companies adopting or dropping a technology, hiring into a function, or entering a geography is how market intelligence teams see movement before it is reported.

Competitive intelligence tools read the same feeds pointed at named rivals: their hiring, their stack, their funding, their web traffic.

Investment analysis and risk

Investment analysis leans on alternative and company data. Funding events, headcount trajectory and technology adoption are the observable proxies for private companies that disclose nothing.

Risk and procurement read ownership chains and registry status instead. Three suppliers that turn out to share a parent is concentration nobody modelled, and it surfaces in company data rather than in a contact database.

Key features to compare

Beyond the data itself, five features separate products inside every category.

Enrichment triggers. Do fields populate on form fill and record creation, or only on a manual run. Automatic enrichment compounds; manual enrichment decays.

Visitor identification. Resolving anonymous website traffic to named companies. Offered by intent and enrichment vendors, and the most qualified first-party signal most teams have available.

Predictive analytics. Scoring accounts by fit or buying stage rather than reporting raw attributes. Useful once there is enough pipeline history to calibrate, and misleading before that.

Free plan or trial. Apollo, Lusha, People Data Labs and Wappalyzer all offer one; the enterprise platforms do not. A free plan is worth more than a demo because it runs on your data.

Data usage rights. What you may do with records after the contract ends, and if the licence covers model training. Data vendors focus their terms here more tightly than buyers expect.

Where LinkedIn Sales Navigator fits

LinkedIn Sales Navigator appears in most of these categories and belongs to none of them. Its data is member-declared and therefore unusually up to date on job titles and headcount, and it exports nothing.

That makes it a research surface rather than a data provider. It cannot feed a segmentation, a sequence or a warehouse, so treat it as a complement to a data purchase rather than as one.

Data security and detailed data

Data security questions get skipped in evaluation and asked in procurement, which is the wrong order. Where records are stored, who at the vendor can query them, and what breach notification looks like are contract terms rather than features.

Detailed data raises the stakes. A dataset carrying named individuals with behavioural history is a different risk profile from a firmographic file, and it should be reviewed as such regardless of which category the vendor sells under.

Keeping data fresh after purchase

Fresh data is a process rather than a purchase, and the decay rates make that concrete.

By the numbers
2.1%

Monthly decay on B2B contact records. Company-level accuracy can deteriorate by 15% monthly across some datasets, and intent signals are worthless after a fortnight. One refresh schedule across all three wastes budget at one end and leaves staleness at the other.

That argues for splitting refresh cadence by volatility rather than running one schedule across everything. Continuous monitoring on contact and event fields, quarterly on firmographics, annual on industry classification.

Enrichment on record creation is the highest-return version, because every account entering the CRM starts complete and the decay clock starts from a better baseline. Retrofitting a stale database costs more and convinces nobody.

How data providers price their products

Three pricing models dominate, and they suit different consumption patterns badly enough that picking wrong costs more than picking the wrong vendor.

Per seat

Standard for sales-facing platforms. You pay for people who log in, usually with a credit allocation attached to cap how much each one exports.

It suits teams where humans read the records and punishes wide distribution. Giving a 60-person sales organisation access to a per-seat platform is where budgets break, which is why most deployments end up narrower than intended.

Per record or per credit

Standard for enrichment and API products. You pay for what you pull, which makes the cost track usage rather than headcount.

Exploratory work is where this gets expensive. An analyst browsing consumes unpredictably; an application enriching every signup consumes predictably. Model both before signing, because credit overage is the most common surprise in this market.

Flat platform fee

Common at the enterprise end and among the dataset vendors. You pay for access to a defined scope, and volume within it is unmetered.

This is the friendliest model for building on, and the hardest to get without an annual commitment. It also removes the internal politics of who gets a seat, which is worth more than it sounds.

What to negotiate beyond price

Credit rollover, seat reduction rights at renewal, a cap on annual increases, and perpetual rights on records already delivered. All four are negotiable at signature and almost none of them afterwards.

Have you noticed?

Vendr's contract data across this category shows average savings against opening quotes running from 12% to 21% depending on vendor. Buyers who settle rollover, seat reduction and price caps upfront pay materially less over three years than those who accept standard paper.

Common mistakes when buying data

Five patterns account for most of the disappointment in this market.

Buying on database size. A provider advertising 100 million records is describing acquisition, not accuracy. The two figures move independently, and only one of them appears in the marketing.

Buying two vendors in one category. Running Crayon and Klue, or ZoomInfo and Apollo, duplicates spend without adding coverage. Running one of each across two categories adds both.

Skipping the compliance check. The only mistake here with legal rather than commercial consequences. Ask how consent was obtained and how DNC screening works before phone numbers change hands.

Not testing on your own list. A week spent checking 100 known accounts settles arguments that otherwise run for a year after signature.

Buying data nobody can activate. If no integration routes the records into the CRM or the sequence, the purchase produces a dashboard. Decide where the data lands before deciding what collects it.

Where to go next

Each category has a detailed ranking with vendors, pricing and the trade-offs between them.

Start with the B2B contact category if you need to reach people, or the company data category if you are building something that consumes records programmatically.

Firmographic and technographic providers cover the attribute layers underneath both. Intent covers timing, and alternative data covers the investor side. Every category above links to its own ranking.