These checks run against your GA4 property on a schedule and can be run again from the dashboard. Each section describes what you see, how we decide the result, and practical limitations.
| Outcome | Meaning |
|---|---|
| Error | Likely hurts reporting or attribution. |
| Warning | Investigate when you can. |
| Info | Informational only. |
| Success | No problem found for this rule. |
| Unknown | Not enough data, or the check does not apply (see message). |
What we look at: For a single reporting day (aligned with GA4 processing lag and your store timezone), we compare paid orders in Shopify to purchase transactions in GA4, after respecting Littledata connection settings that exclude certain orders (for example by source name, app, order tag such as Amazon/eBay, recurring/subscription, or zero-value orders). You may see raw vs adjusted Shopify counts when exclusions apply.
Rules:
| Situation | Typical result |
|---|---|
| Counts match within 2 orders or within 2% | Success |
| Both GA4 and free GA4 connections are active | Error (duplicate purchase tracking risk) |
| GA4 is not connected but counts disagree | Warning (Littledata is not sending to GA4 from this connection) |
| Low volume (both sides under 5 orders) with a gap | Unknown |
| Timezone mismatch could explain the gap | Warning, not Error |
| Middle zone: slightly larger gap than “OK” but not extreme | Warning |
| Large gap (more than 4 orders and more than 5% off), and nothing above explains it | Error |
Optional narrative: When enabled, an AI-generated “Read more” summary may appear in the admin UI with plain-language context. If it is missing, the check still ran—the summary feature may simply be off.
What we look at: Share of sessions in the “Unassigned” session primary channel group (sessionPrimaryChannelGroup — the same dimension as GA4’s Traffic acquisition report), using data from three days ago. We deliberately use session scope, not first-user scope: first-user grouping inherits each returning user’s original acquisition channel, which collapses Unassigned to near-zero and hides real per-session attribution gaps.
We compare that day against your own 30-day average, not a fixed percentage. What counts as normal varies enormously between stores — some run at 15% every day without anything being wrong, while a store that normally sits at 2% has a real problem at 12%.
Rules:
| Outcome | Rule |
|---|---|
| Success | Unassigned is in line with your 30-day average, or a change has not yet lasted two days. |
| Warning | Unassigned is well above your own 30-day average (2 standard deviations, and at least 5 percentage points above it), is at least 15% of sessions, and has stayed there for two days running. |
| Error | Unassigned is over 30% of sessions for two days running — regardless of what is normal for you. |
| Unknown | Fewer than 100 sessions on the day, fewer than 14 days of history to compare against, or GA4 has not finished processing the day (see below). |
Why we wait for two days running: GA4 keeps re-attributing sessions for a day or two after it collects them — paid traffic in particular sits in “Unassigned” until Google matches it to your ad accounts. A single day’s reading is often wrong and corrects itself, so we wait for a change to hold before flagging it. Expect a one-day lag between a real problem starting and the check reporting it.
Why we also check the day is finished: while GA4 is still processing a day, asking it to break that day down by channel can return a large phantom block of “Unassigned” sessions that the same day’s plain session total does not contain — on one store, 56,426 sessions by channel against 40,174 in total, the whole difference sitting in Unassigned. It disappears once the day settles. So we request the day both ways and only judge it when the two agree; if they don’t, the check reports Unknown and tries again the next day. Waiting for a change to hold cannot catch this on its own, because the phantom block appears afresh on every newly-read day rather than passing after one.
Caveats: Often means missing or inconsistent UTMs. Some brands with heavy direct traffic can look “unassigned” more often—that is not always a tagging bug. Because the comparison is against your own history, a level that has been high for a long time reads as normal; the over 30% rule is what still catches those.
What we look at: Whether orders split across marketing channels roughly like sessions do, for the same day (data from two days ago). Large mismatch can mean attribution or tracking issues.
Rules: Uses a standard statistical test. Success when the gap is not statistically significant; Warning for a moderate signal; serious mismatch (Error); Unknown if there are fewer than 40 orders.
Caveats: Real businesses often convert some channels much better than others—so a “difference” is not always wrong. Shown only in internal admin views, not to all users.
What we look at: Monthly event volume vs the 10 million hits nominal free-tier style limit, whether volume is materially above 20M (2× that limit) for a stronger signal, and whether the property is on GA360 (paid).
Rules:
| Outcome | Rule |
|---|---|
| Success | Under 10M/month or on GA360 or (standard) over 10M and under 20M (slightly over free tier). |
| Info | Standard property 20M/month or more — includes indicative GA360 pricing line. |
| Unknown | No traffic estimate available. |
Caveats: The estimate may not exactly match the GA4 UI.
What we look at: Two days ago (settled GA4 day), compared to a 30-day baseline ending four days before today (same window as before). We only flag two situations:
SESSION_TRACKING_SIGNIFICANT_STANDARD_DEVIATIONS in code) and at least 70% below the baseline mean in absolute terms (see SESSION_DROP_MIN_PCT_BELOW_BASELINE). The percentage floor stops very stable stores from being flagged on a “modest” 20–30% dip that happens to be statistically extreme for them — we only want clearly-broken-tracking-style drops. Spikes above the baseline are not flagged.Rules (simplified):
| Outcome | Typical situation |
|---|---|
| Success | Neither rule fired. |
| Warning | Direct-share surge vs typical variation (possible bots / measurement noise). |
| Error | Large session drop or zero from a high-volume baseline (possible removed script / broken tracking). |
| Unknown | No channel history, GA error, or no data for that day. |
You might see: Baseline session mean/SD, aggregate vs daily Direct stats, evaluated calendar date, and flags like sessionDropSuspected / possibleBotTraffic in raw details.
Caveats: Headless stores may be skipped (Unknown). Baseline uses the property timezone to pin the settled calendar day when channel rows omit zero-session days.
Results before 2026-08-27 read low. The channel breakdown behind both rules was requested with an event-scoped channel dimension, which GA4 answers with a small slice of the property rather than all of it — around 3% on a large store. The baseline session average was therefore far below the real one, so rule 1 (broken tracking) could not fire at all, and rule 2 judged Direct share on a fraction of traffic. Both now use the session-scoped dimension. Baselines rebuild over the following 30 days.
What we look at: Same statistical idea as session change, but on transaction (order) counts for two days ago. Direct-traffic heuristics apply to sessions, not this order check.
Rules: Same banding as session tracking for warning vs error when counts collapse to zero vs history.
Caveats: B2B or lumpy order patterns produce more false alarms.
What we look at: Whether the GA4 property has at least one link to a Google Ads account (needed for clean attribution and auto-tagging).
Rules: Success if linked; Warning if not; Unknown if we could not read links.
Caveats: We only see that a link exists, not that conversions or imports are perfect.
What we look at: In the linked Google Ads account, we detect competitor server-side integrations — other providers (for example Elevar or Triple Whale) uploading purchase conversions via the Google Ads API, competing with Littledata’s server-side purchase actions. This is not a flag on “any alternative” conversion action: ordinary client-side actions (website tag or GA4-imported purchases) are expected on most accounts, and when present they are mentioned in the copy so the distinction is clear rather than being flagged. For detected competitor server-side actions we compare attributed conversion value over a rolling window. Littledata’s side of the money comparison focuses on standard “Purchase – Littledata” actions; other action types (for example new customer purchases) may still appear in detection but are not mixed into that dollar matchup the same way.
You might see: Informational outcomes when competitor server-side integrations exist but are secondary; Warning when one is set as a primary conversion action; Error when the competitor out-tracks Purchase – Littledata in a comparable window. Comparisons depend on what Google returns and naming. (This check was previously labelled “Alternative conversions”.)
What we look at: Whether the Google Ads account bids on Littledata’s server-side purchase conversion action. Conversion goals exist at two levels, and this check covers both.
Account level — the account’s default goals, i.e. whether a Littledata conversion action is primary and included in the Conversions metric. This is the main verdict: no Littledata actions at all is a Warning, Littledata actions that exist but are all secondary is a Warning, a Littledata action primary alongside a GA4 purchase action is Info (double-counting risk), and a Littledata action primary on its own is Success.
Campaign level — individual campaigns can replace the account defaults with their own standard goals, a custom conversion goal, or selective optimization. A campaign that simply inherits an account default with Littledata primary is already covered by the verdict above, so it is not reported twice. What this section adds is the opposite case, which the account-level reading alone cannot see: a campaign that overrides the defaults and drops Littledata out, leaving the account looking correctly configured while that campaign bids without Littledata data. Those campaigns escalate the check to Warning.
Where the detail lives: the message stays short — how many of the account’s campaigns Littledata’s goals are active in, then the account-level verdict. The campaigns themselves are named in further info, each one saying which goal it uses (campaign-level goals, a named custom goal, or selective optimization) and which conversion actions it bids on instead of Littledata, so it is clear what is being optimised to.
We also flag campaigns optimising only to “New Customer Purchase – Littledata”. That action is deliberately excluded from every “Purchase – Littledata” value and ROAS aggregation, so a campaign bidding solely on it produces results our reporting cannot see. Campaigns using it alongside the combined purchase action are simply noted, not flagged.
Caveats: Custom conversion goals and selective optimization list their conversion actions explicitly, so those readings are exact. Campaign-level standard goals are inferred from goal biddability plus the action’s primary state. If a campaign declares campaign-level goals but we cannot read them (for example a manager-owned custom goal in a cross-account setup we can’t query), it is left out of the further-info list and never raises a warning — we don’t guess that Littledata is missing. For the same reason a campaign whose goal points at actions we cannot resolve to a name is listed without the “using …” detail rather than with a guess. Cross-account (MCC) custom goals are resolved via the manager where possible.
What we look at: Whether GA4 is linked to BigQuery for export.
Rules: Success if linked; Info if not (with messaging about optional export); Unknown if unreadable.
Caveats: Link existence does not prove exports are healthy.
What we look at: That GA4 has custom dimensions for the eight Littledata parameters Littledata expects (lifetime value, purchase counts, customer IDs, affiliation, store name, app name, etc.). See Littledata’s help article on custom dimensions for customer lifetime value for setup detail.
Rules:
| Outcome | Missing dimensions |
|---|---|
| Success | None missing |
| Info | One or more missing (all treated as informational; not Warning/Error) |
| Unknown | Could not read dimensions |
Caveats: We check presence by parameter name, not pretty names or scopes.
What we look at: How much of your GA4 session traffic comes from undeclared bots. GA4’s own filtering only removes crawlers that announce themselves (the IAB spiders list, matched on user agent); headless browsers and scripted traffic running a real browser engine execute your tags and land in every report. Because the reporting API exposes no individual sessions or IP addresses, we group traffic into signature cells (city × region × channel × OS × screen resolution) over a settled 7-day window ending yesterday (widened to 28 days for stores under ~2,000 weekly sessions, so the share is stable rather than a one-day reading). Every share is computed against the same rows it was counted from, so a cell’s share can never exceed 100% however GA4 splits its rows.
A cell only counts as bot traffic when both of these hold:
800x600, 1280x720, 1280x1200, 1800x1125, (not set), or any square screen such as 1366x1366 — real displays and phone viewports are never square), or its exact fingerprint (city + OS + screen, on unattributed or Direct traffic) matches machine-behaved traffic we have already seen on five or more other stores we monitor — the same actor crawling many merchants leaves the same signature everywhere, which no single store’s data can show.page_view / session_start / first_visit and brought a brand-new visitor for every single session. The second case matters because a crawler that simply waits ten seconds on the page it fetched clears GA4’s engagement timer, so GA4 reports it as engaged; on one store a cell of 52,158 such sessions read as 87% engaged while carrying 3.2 events and a new visitor every time.Neither half alone is evidence: data-centre towns have real residents, dead Direct traffic is what privacy tools look like, and a genuinely new audience skimming one page is not a bot. Sessions reporting no city, no OS and no screen at all are never counted as bots — that fully-blank shape is what GA4 records for visitors who declined analytics cookies (consent-mode pings), whereas a headless bot runs a real browser engine and does report an OS and screen; they are tracked separately in the check’s details and, when sizeable, explained in the check’s detail view. Supporting signals: a single signature owning ≥5% of all traffic, zero-engagement sessions concentrated on one landing page, an all-first-time visitor mix, and the contrast against the rest of the store’s engagement.
Rules: severity tracks whether there is a remedy, not how strong the evidence is.
| Outcome | Rule |
|---|---|
| Success | No signal cleared its floor — no significant cluster of bot-like sessions. |
| Info | A firing verdict (probable: one primary signal; proven: at least two primary plus two supporting). These crawlers carry ordinary identifiers and reach GA4 through the store’s own tag, so nothing Littledata ships removes them — the finding is reported, but there is no action to demand. |
| Warning | A sizeable block (≥5% of sessions) of unidentified events — the consent-ping population above — with Bot Protection still off for the GA4 connection. That is the one state with a fix: turning on Bot Protection stops Littledata forwarding those server-side events. |
| Unknown | Under 500 sessions even in the 28-day window, or the GA4 API could not be queried. |
Benchmark and history: every judged store — clean or not — records its bot-session share, and a firing store’s message compares it against the median across stores judged under the same scorer version (shares from different rule versions are never mixed). Each store also keeps a trailing history of its last ~26 judged shares, so an escalation — a store going from 1% to 60% in a fortnight, which happened in August 2026 — is visible from the check’s own records.
Caveats: Shares are computed from aggregate cells, so counts are detection evidence, not an exact census. The data-centre town list is deliberately conservative (big cloud metros like Frankfurt or Dublin, Ireland are excluded — real audiences swamp them); bot traffic from unlisted hosts still surfaces through the Unassigned, signature and fleet-fingerprint signals, but a low data-centre share is not proof of clean traffic. Both curated lists grow from measurement: each run records the store’s dead traffic by city and its machine-behaved-but-unmatched fingerprints, and entries recurring across many stores get promoted (towns by human review; fingerprints automatically at the five-store bar, constrained to unattributed/Direct traffic with a named city so ordinary bounce traffic can never qualify).
The “Duplicate orders” check was removed: implementation is obsolete (see Shopify vs GA4 order counts above), and its historical auditChecks rows were purged by a one-time script that has since been deleted (see engineering/MIGRATIONS.md). Some accounts may still have historical rows for “Shopping events (funnel)” (and similar retired checks) that no longer run on the daily schedule. Ignore stale labels unless your success team points to a specific historical report.