Back to all datasets
Data listing for Snowflake and Databricks

Digital Personas

Cross-platform behavioural profiles describing how one person behaves across twelve digital surfaces, from browsing and search to social, AI assistants and retail. One row per person, aggregated over a full year.

12
Surfaces per market
670 to 732
Features per profile
US, GB
Markets
Quarterly
Refreshed
12 months
Window

What this is

Twelve surfaces, one person

Most cross-platform data is stitched from separate panels. Here the twelve surfaces sit on the same individual, so breadth, overlap and mix are measurable directly rather than modelled. The median person is active on three of them.

Why it is different

Why it is different

Reach and overlap, not just reach

Because every surface is observed on the same person, you can ask who is present on which combinations, rather than comparing separate panels and hoping they align.

AI assistants included

ChatGPT, Gemini and WhatsApp AI sit alongside search, social and retail, so assistant use is measurable next to the surfaces it competes with.

Coverage stated, not implied

Every per-surface number distinguishes measured activity from an observed person with no events, and from a person not observed there at all. Rates stay honest because the three states are never conflated.

What is in each profile

23 blocks, named by surface

Core activity
Per surface: active days, average daily events, consistency ratio, tenure days, first and last active date.
Temporal
Monthly and daily shares per surface across the full window, plus weekday and time-of-day distribution.
Engagement and mix
Volume and intensity per surface, with a cross-platform summary describing breadth and where a person's activity concentrates.
Category shares
Share-of-activity columns per surface covering event type, content type, retailer and order status. The fifty most variable per surface are kept, so they separate people rather than exhaustively taxonomise.
Purchase
Amazon carries spend, items, unit prices, order status and category. TikTok Shop and Google Shopping are narrower.

How teams use it

What it supports

01

Cross-surface segmentation

Cluster people on breadth, intensity and surface mix. This is the shape the dataset is built for.

02

Reach and overlap

Who is present on which surfaces, and in what combinations.

03

Rhythm

Weekday skew, consistency, burstiness and tenure per surface.

04

Category affinity within a surface

Share columns are directly comparable across people on the same surface.

Specifications

What ships

DatasetDigital Personas, build v1
GrainOne row per person per market, aggregated over the window. Not longitudinal: there is no within-window panel structure and no event sequence.
AggregationDerived behavioural features only: counts, shares, rates and dates. No event logs, no URLs, no free text.
Field count732 columns in the US, 670 in GB. Share columns are selected per build, so intersect the columns before comparing feature by feature.
SurfacesYouTube, Google Search, Instagram, Chrome, TikTok, Google Shopping, ChatGPT, Facebook, Amazon, Gemini, Twitter/X, WhatsApp AI.
Coverage statesA value is measured activity. A 0 means the person is observed on that surface but had no events. A blank means they are not observed there at all. Filter on the surface's total events before averaging any rate.
DemographicsAge band, gender, household income, education, employment and device, on each row and fully populated.
GeographyUnited States, United Kingdom. Each market is aggregated independently, so percentile columns rank within a market and do not compare across them.
Window12 months, June 2025 to May 2026
Refresh cadenceQuarterly
FormatNative tables on Snowflake and Databricks
PIINone. Identifiers are hashed, and there is no name, email or address in the file.

Privacy & provenance

Aggregated features only

Every column is a derived, aggregated behavioural feature computed over twelve months: a count, a share, a rate, a date. Event logs are never exposed and are not part of what ships. All data comes from participants who opted in through the Measure app and browser extension. Participants are compensated, are told what is collected, and can withdraw at any time.

  • No personal identifiers are included. Identifiers are hashed and irreversible.
  • Because the columns are annual aggregates rather than event logs, they carry no timestamps or URLs.
  • The dataset supports no event-level or journey analysis. Sequence is not present.
  • Handled in line with GDPR, CCPA and applicable data protection law.

Request access

Talk to us about scope, sample data and licensing.

Snowflake Marketplace, Databricks Marketplace