Quick definition: Source and medium are campaign-classification fields: source identifies where recorded traffic came from, while medium describes the broad marketing or referral mechanism used to reach a destination.
What are source and medium?
Source and medium are labels used to organize incoming traffic and conversions. A source commonly names the referrer, publisher, platform, partner, or campaign origin, such as a search engine, newsletter, affiliate, or social network. A medium describes the broad mechanism, such as organic search, paid search, email, referral, display, or social. Together they answer a descriptive question: under our documented rules, how was this recorded visit classified?
The labels are not an explanation of a person’s motivation or proof that a channel caused a conversion. A person may see several messages, return directly, use another device, or receive an offline recommendation before purchasing. Source and medium summarize observed acquisition context; they do not establish the counterfactual outcome without the channel. They should therefore be kept separate from causal claims about marketing impact.
Different tools use different defaults for referral handling, direct traffic, paid-search detection, campaign overrides, and session expiration. A source/medium report is meaningful only when the organization documents its taxonomy, precedence rules, and measurement coverage.
Boundaries between source, medium, campaign, and attribution
Source answers “which origin?” medium answers “which general method?” A campaign is a more specific initiative, such as a product launch or audience program. Content or creative fields may distinguish placements or messages. Attribution is the rule that assigns a conversion to one or more qualifying touchpoints. A campaign tag can supply source and medium, while an attribution model decides whether that touchpoint receives conversion credit.
| Field | Example | Question it answers |
|---|---|---|
| Source | newsletter, search-engine, partner-site | Where did recorded traffic originate? |
| Medium | email, organic, referral, paid-social | What broad mechanism was used? |
| Campaign | spring-onboarding | Which coordinated initiative? |
| Attribution model | last click | Which touchpoint receives conversion credit? |
“Direct” deserves caution. It may mean typed navigation or a bookmark, but it can also result from unavailable referrer information, untagged links, app handoffs, browser behavior, privacy settings, or tracking failures. Treat direct as a classification outcome, not an absence of marketing. Similarly, a platform-reported source can differ from first-party logs because identities, windows, and event processing differ.
Measurement and data-quality implications
Create a controlled taxonomy before campaigns launch. Define allowed sources and media, ownership, naming format, capitalization, fallback behavior, and whether each value represents paid, owned, earned, or partner traffic. Validate tags at link creation and normalize values at ingestion. Without this, “Email,” “email,” “newsletter,” and “crm” become separate rows, while an untagged campaign may fall into direct or referral.
Preserve both raw and normalized values. Raw values support debugging; normalized values support stable reporting. Version the classification logic so that a change to channel grouping does not rewrite historical trend lines silently. Record event time, landing URL or sanitized campaign fields, referrer availability, consented-measurement status where relevant, and the rule that produced the final classification. Use metric governance and data validation to assign ownership and test the pipeline.
Limit personal data in tags. Source and medium should describe a campaign, not encode an email address, account ID, phone number, audience membership, search text, or internal customer information. Tags may be copied into browser history, logs, analytics vendors, support screenshots, and referrals. A short controlled vocabulary is more reliable and less risky than flexible free-form parameters.
Experimentation implications
Source and medium can define a pre-treatment acquisition cohort, but only if captured before assignment and used consistently. A landing-page experiment might include only users arriving from a tagged paid-search campaign. Record eligibility at entry, randomize within that population, and compare outcomes across variants. Do not exclude users because a later event changed their source classification or because they converted through a desired channel.
Channel labels can be affected by treatment. A new flow may prompt users to open an email or share a link; a new authentication step may change referrer availability. If the primary outcome is revenue only among visits later labeled “email,” the analysis conditions on behavior that treatment might change. Measure total outcome in the assigned eligible group, then inspect source/medium changes as diagnostic evidence.
For channel investment, attribution reports show where credit was assigned; a randomized holdout or geo test estimates whether the channel created additional outcomes. A campaign can receive little last-click credit and still cause demand, or receive much credit from people who would have converted anyway. See attribution and incrementality for the distinction.
Scenario: paid social landing-page test
A retailer runs paid social ads to a new product page. Campaign links use a standardized source of social-network, medium of paid-social, and campaign of summer-launch. Visitors who meet documented entry rules are assigned to page A or page B before any engagement happens. The exposure event stores variant, time, normalized campaign fields, and an opaque session or account key.
After launch, the team sees fewer purchases classified as paid social for B but higher total purchase completion among assigned visitors. Investigation finds that B asks users to authenticate earlier, and some later visits lose the original campaign context. The team does not claim that B harmed paid social. It reports the pre-specified total conversion outcome, checks assignment and identity coverage, and treats the classification change as a measurement issue to resolve.
The taxonomy owner then updates a test suite for app handoffs and login redirects, documents the new behavior, and marks the date of the rule change. This preserves the usefulness of source/medium reporting without confusing a technical classification change with a marketing effect.
Caveats and common mistakes
- Using source and medium as causal proof. They describe observed classification, not incremental impact.
- Allowing free-form values. Inconsistent names fragment reports and make governance impossible.
- Embedding user data in tags. URLs and event logs are poor places for personal identifiers.
- Changing channel rules without versioning. Historical comparisons become uninterpretable.
- Assuming direct means no prior marketing. Missing referrers and untagged paths are common.
- Filtering experiment results on later channel labels. This can introduce post-treatment bias.
A responsible source/medium workflow
- Define the taxonomy. Publish allowed values, examples, owners, and fallback rules.
- Generate links centrally. Use templates or a controlled builder rather than manual spelling.
- Validate at collection. Detect invalid, missing, or sensitive values before they spread.
- Retain provenance. Store raw values, normalized values, classification version, and event time.
- Monitor coverage. Track direct, unknown, untagged, and rejected traffic by platform and release.
- Separate reports from causal tests. Use stable classification for reporting and credible comparisons for budget decisions.
FAQ
Is source the same as campaign?
No. Source names the recorded origin; campaign identifies a specific initiative. A single source can carry many campaigns.
What should medium values be?
Use a small documented vocabulary that matches reporting decisions, such as email, organic, referral, paid search, and paid social.
Why does direct traffic rise after a product change?
Referrer loss, untagged links, redirects, identity changes, or browser behavior can affect classification without changing demand.
Can source and medium be used as experiment metrics?
They can be diagnostic dimensions, but a primary causal metric should be defined for the full assigned eligible population.
Should we include customer IDs in campaign tags?
No. Use controlled campaign labels and an approved, separate identity mechanism only where a join is necessary.
Summary
Source and medium classify recorded acquisition context: source identifies an origin and medium identifies a broad mechanism. They are essential for consistent traffic reporting but do not prove marketing causality. Govern a limited taxonomy, preserve classification provenance, keep tags free of personal data, monitor direct and unknown coverage, and use randomized incrementality designs for material channel decisions.
Sources
- Google Analytics campaign and traffic-source documentation
- IAB measurement guidelines
- NIST Privacy Framework