← Insights
Insights · Revenue Operations

The week-one CRM data audit, before any AI touches your pipeline.

Every engagement we run starts the same way: two weeks inside the client's CRM, finding out what is true, before a single recommendation. Not because the data is bad. Because nobody knows yet. Here are the six checks, what each one usually finds, and the fix that survives past the next quarter. You can run them yourself in an afternoon.

·8 min read·by zRev·AI-readable edition

Why the audit comes first

A person reading a CRM fills the gaps with memory. The rep knows that "Enterprise" in the segment field means anything over 200 seats, that half the closed-lost deals are still open because nobody closes them, that the lead source says "Website" on everything the SDR team ever touched. A model has no memory to fill with. It takes the data at its word. If industry is blank on half the accounts, it learns that industry does not matter. If lost deals never close, it never learns what losing looks like.

That is why the first two weeks of an engagement are spent finding out what is true, with read access and a spreadsheet, and why the deliverable is a data-quality audit rather than a plan. AI does not create data problems. It stops everyone working around them. Better to know which ones you have before you build a scoring model on top of them.

Six checks. Each takes an hour or two with export access and no special tooling. The order matters less than finishing all six, because the findings compound: an ownership gap looks like a coverage gap until you know which is which.

1. Stages: does the same stage mean the same thing to every rep?

What to pull. For every open opportunity, the current stage, the date it entered that stage, and the rep. Then the written definition of each stage, if one exists.

What we usually find. There is no written definition, or there is one from two years ago that nobody has read. "Proposal" means a deck was sent to one rep and a verbal yes to another. The stage-by-stage conversion rates on the dashboard are an average of several different funnels, which is why they never predict anything.

Why it matters for AI. Stage is the single most informative field for forecasting and for scoring, and it is only informative if it is consistent. A model trained on inconsistent stages learns each rep's habits instead of the buyer's progress.

The fix. Entry and exit criteria for each stage, written in one sentence each, agreed by the reps in the room and enforced by the CRM: a deal cannot enter "Proposal" until a proposal document is attached, cannot enter "Negotiation" until a decision date is set. Discipline lasts a quarter. A required field lasts.

2. Outcomes: is every old deal closed, with a reason?

What to pull. All opportunities with no activity in 60 days, grouped by stage. The list of closed-lost reasons in use, with counts.

What we usually find. A long tail of deals that went quiet and were never closed, so the pipeline number includes money that left months ago. The lost-reason list has 30 options and "Other" is the most popular one, followed by "Price", which reps pick when they do not want to write anything.

Why it matters for AI. Won and lost are the labels. Every scoring model, every forecast, every "which deals look like the ones we win" question is answered from the set of closed deals. If a third of the losses are sitting open at stage two, the model is learning from a sample that is missing its most common outcome.

The fix. Close everything stale in one pass, with the rep in the room so the reasons are real. Cut the lost-reason list to six or seven that a person can actually distinguish. Then automate the stale-deal flag so the pile never grows back: 45 days without activity, the owner gets a nudge; 60, it goes to their manager.

3. Identity: one company, one record?

What to pull. Accounts sorted by normalized name and by domain. Contacts with no account. Deals with no contact.

What we usually find. The same company three times: once from a list import, once from a form fill, once because a rep typed it. Contacts floating with no account because the enrichment tool created them that way. A duplicate rate between 5 and 15 percent is normal in a CRM nobody has audited; it is not a sign of neglect, it is what happens when five systems write to one table.

Why it matters for AI. Routing, scoring and forecasting all assume that an account is one thing. Duplicates split the signal: half the activity on one record, half on the other, and neither looks like a hot account. Routing by territory sends the same company to two reps.

The fix. Merge by domain, not by name, and merge once with a clear survivor rule. Then catch duplicates at the door: match on domain at creation, block or flag, and give the enrichment tool an account before it gets a contact.

4. Coverage: are the fields you intend to score on actually filled in?

What to pull. For the ten fields a scoring or routing model would depend on (industry, employee count, region, tech stack, revenue band, role, source, and whatever your ICP is written in), the fill rate across accounts created in the last 12 months.

What we usually find. The fields that reps have to type are 30 to 50 percent filled, and what is filled is filled inconsistently ("SaaS", "Software", "B2B software"). The fields an enrichment tool populates are 90 percent filled and consistent. The ICP, as written in the sales deck, refers to at least one attribute that does not exist as a field at all.

Why it matters for AI. A model cannot score on what is not there. Worse, a half-filled field is not neutral: the blank half is usually not random (it is the accounts that came in through a channel that skipped the form), so the model learns the channel, not the attribute.

The fix. Choose the use case first, list the fields it depends on, and fix only those. Every one of them gets filled by automation on creation, from an enrichment source, with a pick-list rather than free text. The fields nobody reports on can stay dirty. Cleaning everything is how data projects fail.

5. Ownership: does every account and deal have a live owner?

What to pull. Records owned by users who have left, are deactivated, or are a shared "Sales Queue" account. Accounts with an owner but no activity from that owner in 90 days.

What we usually find. A departed rep who still "owns" 400 accounts because nobody reassigned them. A queue user holding the inbound leads from a form that was set up before the routing rules. Territory rules that reference a region field that check four told you is 40 percent empty.

Why it matters for AI. Every automated action lands on an owner. Routing a hot lead to a person who left is worse than not routing it, because the system reports it as handled. Speed to lead is a routing problem before it is a hiring problem, and the audit is where you find out whether the routing can work at all.

The fix. Reassign in one pass, then write the routing rules against fields that pass the coverage check, and add a catch-all: anything the rules cannot place goes to a named person, with an alert, never to a queue.

6. Source: what does "lead source" actually contain?

What to pull. The distribution of lead source and original source across the last 12 months of contacts, and the same for deals. Where each value is set: by a form, by an integration, by a person.

What we usually find. Two fields that disagree with each other. "Website" on 60 percent of records, covering everything from an organic search to a paid campaign to a partner referral that filled the same form. A source that was renamed in the marketing tool and never in the CRM, so the last six months show up as a new channel.

Why it matters for AI. Source is how a model learns which channels produce deals that close, and it is the field you will use to prove the AI is working. If the field is wrong, both the model and the measurement are wrong in the same direction, which is the kind of error that looks like success.

The fix. One source field, one owner of its value list, set by systems and never by hand. UTM parameters written through to the CRM on every form. The old values mapped once, in writing, so the history is comparable.

What to leave alone. Free-text notes, old activity logs, fields nobody reports on, the custom objects someone built in 2023. The audit is not a cleanup. It is a map of which fields can be trusted for the use case in front of you, so the next two weeks of design build on the trustworthy ones and route around the rest.

What the audit produces

Three things, all written down. A funnel teardown built from the raw records rather than the dashboards: stage-by-stage conversion and velocity, with the leaks marked. The data-quality audit itself: for each of the six checks, what was found, what it blocks, and what the fix is. And a baseline metrics sheet, agreed with the revenue lead, that says what the numbers are today so that "results" in week eight is a comparison and not an impression.

Then, and only then, the design phase starts. The first automation ships inside it, usually against the ugliest manual workflow the audit surfaced. It is almost always one of the fixes above, because the fastest way to make a data problem stay fixed is to make a system responsible for it.

Running it yourself

Everything above needs export access and a spreadsheet, nothing more. Give it an afternoon, one check at a time, and write down the finding and the fix for each. If you get to the end and the fixes are all "remind the reps", the audit has told you something: the CRM is being held together by discipline, and discipline is the one thing that does not scale with an AI on top of it.

Where this fits

These six checks are the first two weeks of every zRev engagement, the Diagnose phase, and the baseline they produce is the only scoreboard we use in week eight. If you would rather have us run them, that is where Revenue Operations and AI Implementation both begin.

© 2026 zRev Solutions
FAQ RSS Privacy & Terms
/* heading semantics (2026-09-12): labels and card titles are real h2/h3, rendered exactly as before */ h2.sec-label { font-size:inherit; font-weight:400; margin-top:0; margin-bottom:0; line-height:inherit; font-family:inherit; }