Skip to main content
The CEO & CMO's audit of product marketing

Stop asserting. Start proving.

For the CEO who wants proof and the CMO who has to produce it: a practical method for auditing whether product marketing actually moves revenue — 19 evidence-gated criteria across five layers, a maturity score, named fix owners, and a board-ready case.

01 · The shift

The audit is already running

Your product marketing is being audited right now, and you did not commission it. The PMM audit used to be a consulting engagement a company bought every few years. Now AI runs it continuously, in the background: every time a buyer asks an answer engine a question, the engine reads your positioning, your documentation, your comparison content, and your reviews — and synthesizes your claims against your competitors' evidence. The scale is not marginal. 73% of B2B buyers use AI tools like ChatGPT and Perplexity in their purchase research, and 94% used a large language model somewhere in their most recent buying journey. The machine grades your work whether or not you look at the grade.

The same shift reaches inside the building. Field-message drift used to be measured on a sampled handful of call recordings; conversation-intelligence AI now grades every call against approved positioning. Enablement analytics expose whether content is used, not just shipped. The evidence that a rigorous audit always needed — and rarely got — is now sitting in the systems your revenue team already runs.

Here is the part that makes this urgent rather than merely interesting: the criteria this audit grades — clear category framing, evidence-backed claims, complete comparison content — are the same inputs answer engines select on when they assemble a recommendation. So the audit is no longer optional hygiene. You are being graded either way; the only choice is whether you see the scorecard first.

02 · The method

One method, nineteen criteria, five layers

The method is a consolidation, not an invention: nineteen criteria drawn from named, established frameworks, arranged into five layers, with one hard rule — no evidence, no score. The contribution is the consolidation and the evidence gate, and the audit runs on two axes. The first is audience: external, buyer-facing work versus internal, org-facing work. The second is layer: Layer 0, Strategic Foundation; Layer 1, External Proof and Usability; Layer 2, Internal Enablement and Adoption; Layer 3, Business Outcome — with Layer 4, Attribution, acting as a gate on everything Layer 3 claims.

The lineage is public and worth naming. Kirkpatrick's four levels of training evaluation give the layers their shape — reaction and learning below, behavior and results above. Phillips' ROI methodology adds the fifth level that becomes the attribution gate. CMMI inspires the 0–4 maturity scale. The Balanced Scorecard inspires scoring perspectives separately instead of averaging them into mush. RACI separates the score from the owner. April Dunford and Geoffrey Moore power the positioning criteria; jobs-to-be-done (Christensen, Ulwick), MEDDIC, Pragmatic Institute, the SiriusDecisions/Forrester revenue waterfall, and Six Sigma's DMAIC — measure before you improve — power the rest. Nothing here is novel except the insistence that it all runs at once, on evidence.

The five layers of the PMM impact audit Strategic Foundation at the base, then External Proof and Usability, Internal Enablement and Adoption, and Business Outcome, with Attribution drawn as a gate capping the outcome layer. Layers 0 and 1 are what AI reads; Layer 2 is what AI measures. Layer 0 · Strategic Foundation Layer 1 · External Proof & Usability Layer 2 · Internal Enablement & Adoption Layer 3 · Business Outcome Layer 4 · Attribution — the gate no mechanism → outcome score capped what AI measures what AI reads nineteen criteria · evidence-gated · scored 0–4
The stack: foundation and proof at the base, enablement above, outcome on top — and attribution as the cap that decides whether the outcome counts.
03 · The external layers

Layers 0–1: the layers AI reads

The external layers grade whether a stranger — human or machine — can understand what you are, why you win, and what the proof is, without a call. Layer 0, the foundation, has three criteria. (1) Category framing and context: does the message answer "what category is this, and what budget line does it displace" without effort — Dunford's positioning discipline and Moore's category logic, applied as a pass/fail reading of your own homepage. (2) The value and impact chain: an unbroken lineage from feature to capability to benefit to quantified impact, with no missing links. (3) Defensible differentiation: moats — claims a competitor cannot copy in a quarter — rather than feature-parity assertions that read identically on both websites.

Layer 1 grades the proof. (4) Direct competitive clarity: self-serve comparisons that neutralize objections before the first call. (5) Tactile and interactive proof: tours, calculators, sandboxes — evidence a buyer can operate, not just read. (6) The persona/JTBD dual layer: content that resolves the economic buyer's strategic job and the end user's functional job at the same time. (7) Feature-to-value translation completeness: no spec ships without a "so what."

Now the AI turn. These seven criteria are precisely what answer engines evaluate when they assemble a shortlist. Category framing is how the engine classifies you. Evidence-backed value claims are what it quotes — the Princeton and Allen Institute GEO research found targeted optimization lifts generative-answer visibility by up to 40%, and the strongest levers are quotations (+27.8%), statistics (+25.9%), and cited sources (+24.9%). Comparison content is what shortlist queries retrieve. And the engine synthesizes your claims against competitors' documentation and reviewers' experience, so overclaiming gets contradicted inside the very answer that cites you. A low score on Layers 0–1 is no longer a messaging problem; it is a distribution problem. The mechanics of being the answer are the subject of the companion piece on AI search.

04 · The internal layer

Layer 2: the layer AI measures

Layer 2 grades whether the work changes what the org actually does. (8) Discovery-to-recommendation guidance: MEDDIC-style mapping from discovery answers to the recommended solution, embedded where reps work — not a slide in a portal. (9) Aha-moment inventory and demo choreography: the moments that convert, catalogued and staged, not left to each seller's instinct. (10) Message consistency in the field: is field language matching approved positioning, or has it drifted. (11) Content findability and utilization: whether reps open the asset, not whether it exists. (12) Cross-functional pull-through: does PMM's work change roadmap, pricing, and packaging decisions upstream, or does it only decorate what was already decided.

Two of these criteria have been transformed by tooling. Message consistency was once a sampled handful of call recordings graded by hand; conversation-intelligence AI now grades every call. Utilization was once a guess; it is native in enablement analytics. Layer 2 used to be the softest part of any PMM audit. It is now the hardest-evidenced.

Score the outcome. Tag the owner.

A low score on criteria 10 or 11 is usually not a PMM failure. PMM owns the message and the asset; sales ops and enablement own enforcement and surfacing. The fix is not to soften the score — it is to score the outcome and tag the owner separately, which is what RACI is for. This one design rule ends the two ways these audits die politically: PMM graded on levers it does not hold, or real gaps hidden to protect a function. The scorecard stays honest because honesty stops being expensive.

05 · The outcome gate

Layer 3 and the gate: no attribution, no credit

Layer 3 is the layer everyone wants to talk about, and it comes last on purpose. Its criteria are the outcomes positioning is supposed to move: win-rate impact; competitive win-rate delta against the named competitors the positioning targeted; sales-cycle compression at the targeted stall point; ACV and deal-size impact; stage-conversion velocity at the targeted funnel stage; category and analyst perception shift; and new-rep time to productivity.

Then the gate, from Phillips' Level 5: does a mechanism exist that isolates the initiative's effect, or is the connection asserted? The mechanisms, in descending rigor: a control or holdout rollout — ship the new message to part of the sales org, hold a comparison cohort back, the incrementality testing marketing measurement already uses; trend-line analysis adjusted for confounds; disclosed estimation, where the assumptions are written down and challengeable; and win-loss tagging of asset usage per deal — supporting evidence, though tagging alone does not open the gate. The design rule is stated plainly: no attribution mechanism means outcome claims are discounted and the outcome score is capped at 2 of 4, regardless of how good the raw number looks. A great number with no mechanism is a coincidence wearing a suit.

One modern note. Holdout rollouts used to be politically and logistically expensive — nobody wanted to be the cohort that kept the old deck. With AI drafting the variant assets and CRM automation splitting the cohorts, a 30-day holdout is now a realistic ask, not a research project.

06 · The score

The score, and what it buys you

Each criterion scores 0–4 on a CMMI-style maturity scale: 0, absent; 1, ad hoc; 2, defined; 3, managed; 4, optimized and tied into attribution. Layers weight into the composite at 15 parts foundation to 25 parts each for external proof, enablement, and outcome; attribution carries no weight of its own — it acts as the cap. The composite lands in one of four bands: Asserted (below 1.0), Documented (1.0–1.9), Measured (2.0–2.9), Proven (3.0–4.0). The band names are the thesis: most product marketing lives in Asserted, where every claim is sincere and none is checkable.

What the leader does with the score: strategize — the two-axis grid shows at a glance where the strategy is thin, external or internal, foundation or proof. Find gaps — each low score arrives with its evidence and its owner attached, so the fix list writes itself. Progress — re-score quarterly; the composite delta is the operating metric, the one number that says whether the function is compounding. And sell it upstairs — which deserves its own section.

07 · The board case

Selling it upstairs: the audit as a budget instrument

Boards discount marketing self-assessment because it is usually unfalsifiable — every function's deck says the function is working. An evidence-gated score with named owners and an attribution cap is the opposite of unfalsifiable: it can lower its own grades, which is exactly what makes its high grades worth funding. The asks arrive pre-structured — each gap maps to an action, an owner, a cost class, and the specific Layer 3 metric it should move. And the AI-visibility angle converts gaps into urgency a board already understands: the same failures that flatten win rates also erase you from AI answers, and the buyers who do arrive from AI answers are the ones worth having — AI-referred visitors convert at 14.2%, about five times Google organic.

The RACI tags do the final piece of work: the ask is not "more budget for marketing." It is a cross-functional fix list with the right owner on each line — some lines belong to PMM, some to enablement, some to RevOps. A board can fund that, because it reads like an operating plan rather than a plea.

And if you sit in the CEO seat, run the same instrument in reverse. Do not ask your marketing leader whether product marketing is working — every sincere answer to that question is a narrative. Ask for the scorecard: nineteen scores with evidence attached, an owner on every gap, and the attribution cap intact. The request takes one sentence, the first honest version takes thirty days, and it replaces the least falsifiable slide in your board deck with the most falsifiable one.

The board-ask format — three rows shown (invented example)
GapOwnerAskMetric it moves
Comparison content is thin; competitive claims carry no evidencePMMOne writer-quarter to build self-serve comparison pages with cited proofCompetitive win-rate delta vs. named competitors
Field language has drifted from approved positioningEnablement (PMM supplies the message)Conversation-intelligence scoring of approved language on every callMessage-consistency score; stage conversion at discovery
No attribution mechanism on the current launchRevOpsCRM cohort split for a 30-day messaging holdoutIsolated win-rate impact of the launch
08 · Start here

Start here. Thirty days.

Days 1–5: gather the evidence and take the honest first score. The evidence classes: your positioning and product pages, one battlecard or comparison asset, a sample of call recordings or win-loss notes, enablement usage analytics, CRM funnel metrics, and a fresh-session AI answer check on your category — ask the engines what a buyer would ask, and read what comes back.

Days 6–15: score the nineteen and publish the scorecard internally. Owners tagged on every gap, evidence attached to every score, no diplomatic rounding. The first scorecard is always the worst one; publish it anyway.

Days 16–25: stand up the attribution mechanism on one in-flight initiative. A holdout if you can get it; disclosed estimation if you cannot. One initiative, instrumented properly, beats a portfolio of asserted wins.

Days 26–30: build the board case from the scorecard. Gap, action, owner, cost class, metric — the table above is the whole format.

pmm-impact-audit-skill

The skill that audits with you

Paste your evidence — it scores what it can see, refuses to score what it cannot verify, tags the fix owner on every gap, and hands back the scorecard, the roadmap, and the board case.

Close

Get audited on purpose

The alternative to this audit is not "no audit." It is being audited only by machines, continuously, with no chance to fix what they find. This site practices the discipline it prescribes: the live fit-checker on the homepage answers only from a vetted fact base — claims gated by evidence is not a theory here, it is the architecture.

This page is built to be forwarded. From the CEO seat, send it to your marketing leader with one sentence: "I want the scorecard." From the CMO seat, send it to your product marketing team with "score us" — and take the result to the board yourself.

The skill is free, like everything else I publish. If you want this capability built in your organization rather than just read about — leave your email below, or reach me at daniel@cmoconfessions.com. No gate, no sequence.

Optional — only if you want to talk. The skill above needs nothing from you.