Playbook · Attribution

The Attribution playbook.

Everything an architect or engineer needs to run one of these projects end-to-end — Blueprint → Build → Enable → Maintain. The most maintenance-heavy motion we run, and the one that finally lets a marketing team prove where its pipeline comes from. Start with the five-minute audio brief, then open the repo.

4Phases
2Questions · what & where
3Touches · First · MQL · Latest
SF + HSBoth platforms
Watch or listen first · 5 min

The audio brief.

One narrator, five minutes, the whole shape of the project — now with the deck that runs alongside it. Play it before you open anything else; it's the fastest way to get into the right mindset for an attribution build. In a hurry? Change the speed under the deck — it sticks next time.

0:00 / 0:00
Speed
1 / 1
Chapter 1 Narrated brief · generated with ElevenLabs · Download ↓
Chapters
The whole idea

Attribution answers two questions, not one — what did they do (lead source) and where did they come from (channel). Nail that split, tag every touch the moment it lands, and you can finally say where the next dollar should go.

The Shape

Four phases, one motion.

Every Attribution project moves through the same four phases. Know what you produce in each — and know that, like CPQ, this is a project you win or lose in the Blueprint, then keep winning in Maintain.

Start Here

Kick off a project in one paste.

Don't dig through the GitHub repo. Copy this prompt, drop in the customer name, and send it to your Claude. It clones the template, reads the AGENTS.md, and walks the Blueprint checklist with you.

paste into Claude
# Attribution — kick off
Clone the LeanScale Attribution template repo for my customer [Customer Name].

Then read AGENTS.md and run the Blueprint checklist:
  1. Confirm the questions the exec team wants this attribution model to answer (this sets the depth).
  2. Map where data enters the CRM today — web forms, lead-list uploads, rep-created contacts, event lists, partner regs.
  3. Map their world onto our standard taxonomy — lead sources (what) and channels (where), with AI/AEO sources included.
  4. Decide the model — First / MQL / Latest touch, and how multi-touch gets tracked (campaigns).
  5. Produce the tagging map — every source & channel, tagged programmatically or by a written manual process.

Ask me anything you need before you start.

What you get

A customer-specific repo

The clone becomes part of the customer's brain — the taxonomy, the field map, and every tagging decision from this project inform their agent for the next one.

Inside it

The standard taxonomy + a field package

Our lead-source and channel taxonomy as a starting point, the HubSpot & Salesforce field packages, the UTM→channel translator, and the tracking-code snippet.

1
Phase 1

Blueprint.

This is the whole game. Before you touch a field, get the exec team to tell you what decisions this model has to drive — that sizes everything. Then map how data actually enters their system today, and lay it onto our standard taxonomy. The projects that go sideways are the ones that got the depth wrong (too much granularity nobody uses, or too little to answer the question) or where a source was defined but no mechanism ever routed it.

1.1 · The Two Questions

Lead source is the what. Channel is the where.

This is the one thing to internalize before anything else. Attribution isn't a single field — it answers two different questions with two different sets of fields, and people conflate them constantly.

Lead Source — what did they do?
  • The action that got them into the CRM
  • Registered for a webinar · requested a demo · attended an event · accepted a gift
  • Answers: which content actually converts?
  • Drives: do we do more of this?
Channel — where did they come from?
  • The path that drove them to take that action
  • LinkedIn ad · sales outbound · organic search · partner · a generative-engine referral
  • Answers: where should the next dollar go?
  • Drives: which platform earns more spend?

The house analogy

Think of designing a house. The lead source is the garage — the room they ended up in, the thing they did. The channel is the door they walked in through — where they came from to get there. Same room, different doors. That's why they're two fields: one answers what content works, the other answers how to get more people to it.

Worked example · one webinar sign-up
Contact · Lead
Channel · whereLinkedIn Ad Channel · orSales Invite Channel · orDirect Lead source · whatHosted Webinar

Three people register for the same webinar (one lead source) from three different channels. Lead source tells you the webinar converts; channel tells you the LinkedIn ad — not the cold outbound — is what to fund next.

1.2 · Three Levels of Granularity

Each question nests — group → item → detail → exact.

Both dimensions are a set of Russian dolls: a high-level bucket, the value itself, a level of specificity, and a free-text exact. Depth is a choice — match it to the questions, don't max it out by reflex.

Lead source · the what, nested
  • Grouping — Event (exec-level bucket)
  • Lead Source — Hosted Webinar (the action)
  • Lead Source Detail — Registered / Attended / No-Show
  • Lead Source Exact — "State of GTM · 2026" (free text)
Channel · the where, nested
  • Channel Grouping — Paid Digital (exec-level bucket)
  • Channel — Paid Social (the path)
  • Subchannel — LinkedIn (the platform)
  • Campaign — the UTM campaign name (the specific initiative)

Rule of thumb

Only add the grouping layer once you're past ~15 lead sources or ~10–15 channels — before that it's overhead. And restrict source, channel, and subchannel to an agreed pick list, never free text below the exact field. A hundred spellings of "LinkedIn" makes the whole report undecidable.

1.3 · The Standard Taxonomy

Don't start cold. Start from our standard, then tailor.

Most marketers have some attribution experience but haven't gone deep. Hand them our taxonomy as the starting point and edit from there — it's far faster than a blank page. Here's the shape of it.

Lead sources · the what
Groupings
GroupingWebsite GroupingEvent GroupingEmail GroupingReferral GroupingGifting

Under Website: Demo Request, Community Signup, Website Content (guide / on-demand webinar / blog), Chatbot. Under Event: Hosted & Third-Party Webinar and In-Person Event, each with a detail of Registered / Attended / No-Show / Visited Booth. Under Referral: Partner, Customer, Employee, Personal Contact.

Channels · the where
Groupings
GroupingOrganic Digital GroupingPaid Digital GroupingReferral GroupingEvent GroupingEmail

Organic Digital: Direct, Organic Search (Google/Bing), Organic Social (LinkedIn/YouTube/…), Review Sites. Paid Digital: Paid Search, Paid Social, Paid Review (G2). Referral: Web, Partner, Customer, Employee, Sponsored & PR content — and Generative-Engine Referral. Email: HubSpot, Marketing Outbound, Sales Outbound.

One caveat on "Event"

Event lives in both dimensions, and that's the tell for the whole model. It's a lead source (they attended) and a channel (offline). The useful question is rarely "did they come from an event" — it's what drove them to register: a sales invite, a paid ad, an in-house email. Capture that on the channel + campaign, and the event ROI question answers itself.

1.4 · AI & AEO Sources

Add AI and AEO to the taxonomy — every time.

Our old attribution standard was written for a pre-AI world and is missing generative engines entirely. In 2026 that's a hole in the funnel. Answer Engine Optimization traffic is real pipeline, and if it isn't in the taxonomy it reads as direct. Make it first-class.

Generative-Engine Referral
New channel · subchannel
  • A dedicated Referral › Generative-Engine Referral subchannel
  • Last-referring sites: chatgpt.com, perplexity.ai, gemini.google.com, claude.ai
  • UTM medium generative-engine when you can tag it
"Other" organic search
Catch the assistants
  • Search subchannel "Other" catches ChatGPT, DuckDuckGo, Yahoo as referrers
  • Where the tracking code reads the last referring site
  • Keeps AI answers out of the "direct" bucket
Report it separately
So it earns a budget
  • Once it's a value, it shows up in the channel roll-up
  • Now AEO can be compared against paid for cost-per-pipeline
  • The refresh we fold back into the standard on every build
1.5 · Where Data Enters Today

Start from the entry points, not a blank taxonomy.

The fastest way in: give the client homework — come to the kickoff with a flowchart of how records enter the system and where they go. That map hands you the lead sources and channels to account for, and the inflection points where you need to answer "where did this come from."

1
List every entry point
Web forms (HubSpot), lead-list uploads (post-event), onesie-twosie rep creation, LinkedIn / ZoomInfo imports, partner registration forms, chatbot. Each one is a place attribution can be captured — or lost.
2
Define the high-value actions that mean MQL
A limited set of actions that should hand a lead to sales. This list changes over time and must be owned by someone — if you add a source but no mechanism routes it, it silently never reaches sales.
3
Right-size the depth to the questions
First / Latest touch is the floor; MQL touch is optional. Don't over-engineer — some teams just need "are events producing pipeline?" (LeanScale itself mostly runs on lead source alone). Depth is a decision, not a default.

The Harbinger lesson

Attribution setups morph as the business changes — and break when the client doesn't tell you. Harbinger changed what sales expected to receive but never told us the model needed to change, so leads routed to the wrong place. Bake this into the blueprint: who owns changes, and how do they reach us.

The Deliverable

What Blueprint hands over.

Three artifacts come out of Blueprint, and everything in Build is downstream of them. Get them signed off before you automate a single tag.

01

The taxonomy

The tailored lead-source and channel lists — groupings, values, details, and the restricted pick lists — with AI/AEO sources included. The agreed vocabulary everything else maps to.

02

The tagging map

Every source & channel with a column for how it gets tagged — programmatic (form → workflow) or manual (lead-list upload) — and how it lands in a campaign. The manual ones need a written process.

03

The model decision

First / MQL / Latest depth, how multi-touch is captured (campaigns), opportunity + partner rules — and, above it all, the questions the exec team signed off on.

Use it

Feed all three to Claude as the reference for what Build produces. Every engagement's version looks a little different — but the shape of the deliverable is always these three.

2
Phase 2

Build.

Build has two halves, and the order matters. First you capture the data — because a touch that isn't tagged doesn't exist. Then you stamp it into the fields and campaigns that make it reportable. Get capture wrong and the prettiest field model in the world reports on nothing.

2.1 · The Capture Backbone

UTMs on everything — from a restricted pick list.

This is the backbone of the whole model. Every marketing link, campaign, and channel carries UTMs, drawn from an agreed set of values so the data comes in clean. Give the team a UTM builder so they can't freelance.

A UTM builder app
  • Restricts source / medium / campaign to accepted values
  • Izzy built one on Pat Lytics; Acton shipped a builder + channel mapper
  • Bundle one into this playbook — every attribution engagement gets it
  • Kills the 100-variations problem at the source
The UTM → Channel translator
  • Maps utm_medium + utm_source + last-referring-site → channel & subchannel
  • paid-social + linkedin → Paid Social · LinkedIn
  • Handles non-digital (events, offline referrals) too
  • The one hard-to-maintain piece — keep it in the repo, versioned
2.2 · The Tracking Code

Turn "fake direct" into real attribution.

The move that gets you close to a 100% model. When a visitor lands with no UTMs, empty data looks like direct traffic — but they came from somewhere. A small piece of custom JavaScript recovers it.

1
Read the last referring site
On landing with no UTMs, the script checks the previous session / last referring page (via cookie tracking) and writes utm_source and utm_medium so the form submission captures where they actually came from.
2
Backfill first, then enable going forward
Clean up the historical gap (e.g. an event vendor page that pointed straight at the site with no UTMs), then switch the code on so anything similar is auto-tagged in the future.
3
Know the ceiling — it relies on cookies
If a visitor denies cookies, the trail breaks — it's not foolproof. It's standard web-dev best practice, lightly customized to the client's field mapping. It gets you closer to 100%, not all the way.

Provenance

This is the advanced scripting Sean built — it's what let us untangle real direct traffic from fake direct on the Harbinger build. Keep the snippet and its field-mapping notes in the repo; it's one of the highest-leverage assets in the whole playbook.

2.3 · The Field Package

First touch, MQL touch, Latest touch.

Three touch-sets on the contact, each stamping the full picture — lead source (with detail & exact), channel, subchannel, campaign, and the raw UTMs. The template ships the whole field list for both platforms.

HubSpot — contact object
  • Original · Lead Source / Detail / Exact · Channel / Subchannel / Campaign · UTMs
  • MQL · same set, created on the Deal object
  • Latest · same set + "Latest Attribution Completed" datetime
  • Latest fields are pick lists; original & MQL are free text stamped at the moment
Salesforce — same shape
  • Lead/Contact carry the same First / MQL / Latest field sets
  • Push where there's API surface; otherwise workflow-stamp
  • Opportunity gets its own attribution fields (see 2.5)
  • Campaigns + campaign members carry the multi-touch (see 2.4)
1
Latest touch is the living value — lock it down
Every action with a lead-source component updates Latest via workflow. This is the field that must be a restricted pick list, because everything downstream — MQL, opportunity — stamps off it.
2
Original gets stamped once, when the record is new
If the record was created recently (Spy Cloud uses a 2-hour window), Latest changed, and Original is empty → write all Latest values into Original. After that it's frozen. First touch, captured.
3
MQL stamps off Latest at the MQL moment
When the person hits the MQL stage, copy their current Latest values onto the MQL fields. It's the least necessary of the three — it usually equals Latest anyway — so drop it if the team won't use it.
2.4 · Multi-Touch via Campaigns

First and Latest are two touches. Campaigns hold the rest.

First and Latest capture the ends. For everything in between, use campaigns — the standard way to hold the whole multi-touch lifecycle. It's not weighted; it answers "how many things did they interact with before converting."

1
Add to a campaign on every meaningful interaction
Clicked a LinkedIn ad, registered for a webinar, opened a sequence — each adds them to the matching campaign, automatically where you can.
2
Stamp the campaign-member fields at join time
Channel, subchannel, campaign, and UTMs get written onto the campaign member at the moment they join — a snapshot of where that touch came from. Simplify by creating campaigns per lead source and carrying channel/UTMs on the member.
3
Report off membership + membership date
Pull every campaign a contact belongs to, ordered by membership date, and you have the full touch timeline. HubSpot retains some of this natively; in Salesforce, campaign membership is your multi-touch record.
2.5 · Opportunity & Partner Attribution

Roll it up to the deal — from the primary contact.

Lead attribution feeds opportunity attribution. Keep it simple: attribute the opp off one person, and treat partners as a channel with a little extra plumbing.

Opportunity attribution
Stamp from the primary contact
When the opp is created, copy Original + Latest source and channel from the primary contact role onto the opportunity — plus a campaign lookup so campaign-influenced reporting works. You can append every contact's values, but it gets convoluted fast; primary-contact is the standard.
Partner attribution
A channel, plus deal-reg plumbing
Partners sit on channel / subchannel (Partner Referral). Set partners up as accounts with deal-reg objects and unique tracking links; forms map back via lookup and auto-add to campaigns. Decide the credit rule up front: net-new leads only, or also deal-regs on opps that already exist? Tools like PartnerStack / Allbound feed this.

Cutover — like every build

Ship quick fixes (a bad UTM mapping, a mis-tagged source) right away. But deep changes — new workflows, the tracking code, field-model changes — go in after hours, when reps aren't in the system (Friday afternoon / weekend), with the broader team briefed the Wednesday before.

3
Phase 3

Enable.

Attribution is the motion that dies fastest without enablement, because it depends on human behavior every single day. The one message that has to land: this is not set-it-and-forget-it. Every new source, campaign, or channel needs UTMs and a home in the taxonomy — or you lose visibility.

📄 Documentation
Read
  • The taxonomy + the tagging map
  • How each source gets tagged, and by whom
  • What "high-value action = MQL" means and who changes it
🎧 Audio brief
Listen
  • Every enablement gets a brief (ElevenLabs)
  • People retain audio they'd skim as text
  • This page's brief is the template
🗂️ The change log
Maintain
  • A living list of every active value that's live out there
  • Deprecate a value → find the links that still use it
  • The single artifact that keeps the taxonomy from rotting
3.1 · The Manual Motions

The human element is where attribution leaks.

Not everything can be automated. A rep who meets someone at a booth and adds them from memory looks like sales outbound forever. The manual paths need a written process and real buy-in on why it matters.

Give every manual path a template
  • A lead-list upload template with the required columns — "I talked to this person at this event"
  • The rule for rep-created contacts (source, channel, campaign)
  • The partner registration form → CRM mapping
Sell the why, not just the how
  • Untracked event leads make it harder to justify the event next year
  • Bad tags don't just miss data — they misroute leads
  • The team has to want clean attribution, or it won't happen
3.2 · Hyper-care Cadence

Brief before, cut over quiet, then office hours.

Same shape as every LeanScale cutover — pre-brief the team, flip it when the system's quiet, then hold office hours while the first real touches flow through and surface what's missing.

Wed · before
Pre-brief
Walk the marketing + sales teams through what's changing and what they now own.
Fri · 2pm
Cutover
Flip the workflows & tracking code when reps are logging off.
Mon · after
Office hours
Watch the first live touches tag correctly; fix mappings in real time.
Fri · +1 week
Office hours
Confirm a full week of data is clean, then hand the change log to the owner.
3.3 · Clone the repo into their brain

We clone this attribution repo for each customer. The taxonomy, the field map, the tracking-code notes, and the tagging decisions all become part of their agent's context — so the next project on that account, and every ad-hoc tweak, starts faster.

Name the owner on day one

Attribution without an owner rots in a quarter. Someone on the client side has to own the taxonomy, the change log, and the "is this a new MQL action?" call.

Enablement isn't a handoff; it's compounding context. Teach the owner enough to add a source and update the translator without us — and to know when to call us back.

4
Phase 4

Maintain.

Be honest with the client: attribution generates more ongoing work than almost anything we run. This is why mops teams look flat and bloated — some companies have a whole role just keeping paid-social attribution correct. Not strategy. Just keeping it true. Maintain is about making that cheap and, wherever possible, the client's job.

4.1 · Ad-hoc Triggers

What kicks off maintenance.

Any one of these adds or changes a value — and every value has to reach the taxonomy, the translator, the workflows, and the change log.

Trigger

A new channel or platform

A new ad platform or social network — needs a subchannel, a UTM convention, and a translator row.

Trigger

A new event or webinar

New exact values and a campaign, plus UTMs on every registration link so it doesn't read as direct.

Trigger

A new lead source

A new form or motion — decide if it's an MQL action, and build the mechanism that actually routes it.

Trigger

A deprecated value

Retiring a source or channel? Use the change log to find every live link that still uses it and update them.

Trigger

A changed MQL definition

Sales changes what they'll accept — the routing and the MQL stamp both have to move with it.

Trigger

A new partner program

New partner accounts, deal-reg objects, tracking links, and the credit rule for existing vs net-new deals.

4.2 · The Drift & the Gotchas

The ways attribution quietly goes wrong.

None of these throw an error. They just slowly poison the report until nobody trusts it — which is why the change log and a periodic audit exist.

Fake direct traffic
The #1 leak
  • Empty UTMs read as "direct" — hiding the real source
  • The tracking code is the antidote; verify it's still firing
Untracked event leads
Human element
  • Reps adding contacts from memory → looks like outbound
  • Enforce the lead-list template after every event
Cookie denial
Structural ceiling
  • Consent-declined sessions break the trail
  • Accept it — the model is "close to 100," not perfect
Value sprawl
Report-killer
  • Free-text creep → 100 spellings of one channel
  • Pick lists + the UTM builder keep it contained
Stale links
Silent drift
  • A deprecated value still live in an old link
  • The change log is how you catch it
Silent expectation change
Relationship
  • The business evolves; nobody tells the implementer
  • Owner + a standing check-in keep you in sync
4.3 · When to Revisit

On events, and on a clock.

Some maintenance is reactive; some has to be scheduled or it never happens. Do both.

Event-based
Revisit when the model changes
  • New channel, event, source, or partner program
  • A changed MQL definition or routing rule
  • A new marketing motion the taxonomy doesn't cover
Time-based
Audit the values quarterly
On a clock, reconcile the change log against what's actually live — dedupe values that drifted, retire what's dead, confirm the tracking code and translator still fire, and re-confirm the questions the exec team is answering haven't moved. This is the work that keeps the report trustworthy.
What You Hand Over

The assets, in one place.

The team gets this landing page with the brief, a one-paste prompt, and the template repo behind it. The client gets a taxonomy, a wired field model, a UTM builder, the tracking code, and an owner who keeps it true.

For the team

This landing page + brief

The playbook overview and the 5-minute audio brief — the front door for anyone running an attribution project.

For the team

The one-paste prompt

Clone-for-customer → AGENTS.md drives the Blueprint checklist, starting with the exec questions and the entry-point map.

For the client

Taxonomy + field package

The tailored lead-source & channel lists, the HubSpot / Salesforce field sets, and the tagging map — signed off before Build, inherited after.

In the repo

The UTM builder, translator & tracking code

The restricted UTM builder, the UTM→channel translator, and Sean's tracking-code snippet with its field-mapping notes — the compounding assets.

The bet

Attribution is our most maintenance-heavy motion — and that's exactly why a repeatable, opinionated playbook is the edge. Anchor on the two questions, tag every touch at the source, include AI/AEO, and enable the client so well they keep it true themselves. This gets us ~60% of the way — the rest we learn on every close-out and fold back in.

Closeout

Debrief the project.

When a Attribution engagement wraps, spend sixteen minutes with the debrief agent. Teamwork already knows what got built and when. This is for the part none of our systems can see — the call that could have gone either way, the thing the customer wanted that you refused, the near-miss that never became an incident, and above all the places this playbook turned out to be wrong. What comes out of it gets written back into this page.

Before you start

Two minutes of thinking beats sixteen minutes of recall. Have these in your head — you don't need notes, and you definitely don't need a script.

  • The taxonomy you actually shipped — how far it drifted from our standard, and what this customer's market forced you to add.
  • The entry points — how many they really had, and which one you nearly missed.
  • AI and AEO — whether those values are live and receiving data, or sitting empty waiting for someone to wire them up.

Sixteen minutes, one sitting. It's a voice conversation, so your browser will ask for microphone access — use headphones or the agent will hear itself. Chrome is the safer bet over Safari.

Hand over the artifacts too

The debrief asks what you built that's worth reusing. This is where you actually hand it over — dashboards, field maps, flows, templates, scripts, spec docs. Files land in the project's Drive folder; links go into the artifact register so the next person can find them.

Thirty seconds, and do it while the project is still in your head. “I'll upload it later” is exactly how assets end up trapped on an account.

Add an artifact →