🎉 Measure Predict is now live 🎉 Agentic cross-platform behavioral intelligence in your pocket. Try for free →
Back to Privacy Notice

How we scrub and anonymise data, free text and journeys

Last updated August 7, 2026

When you take part in a data task, some of what you share is in your own words: searches you typed, pages you browsed, conversations you had with an AI assistant. Before any of that leaves Measure, we put it through a process we call scrubbing.

This page explains what scrubbing does, what it cannot do, and where we have drawn the lines.

The short version

  • We remove the things in your data that could lead back to you. Not just your name, but combinations of details, and specific phrases that could be matched against records someone else holds.
  • Removing your name is not enough, and we do not pretend it is. People can be identified without one.
  • Where something cannot be cleaned up without destroying what makes it useful for research, it does not leave. It stays inside our secure environment, and only summaries, patterns or models come out.
  • Scrubbed data is anonymous to the person receiving it. It is not anonymous to us. We still hold the key that links it to you, which is why you keep all your rights over it.
  • None of this is used for advertising unless you separately opt in.

What is scrubbing?

Scrubbing means going through data before it is shared outside Measure and removing anything that forms a link back to a real person.

We look for what we call bridges. A bridge is anything in the data that could carry someone from the content back to you. Our job is to find every bridge and remove it, using the lightest change that works, so that the data is protected and the research still means something.

Why isn't removing my name enough?

Because people can be identified without a name.

Say a piece of text contains no name, no email and no account, but does say “as a 34-year-old nurse in Brighton with two kids…”. Nobody is named. That combination may still narrow to a very small number of real people, possibly one. It is identifying, and the law treats it that way.

So the test is not “did we delete the obvious identifiers”. It is three harder questions:

  • Can someone be singled out? Is any record specific enough that only one person could match it?
  • Can it be linked? Could this be joined to another record about the same person?
  • Can something be inferred? Could someone work out who this is, or learn a new fact about an identified person?

If the answer to any of those is yes, the data still counts as personal and we have not finished.

We apply the standard the UK regulator expects. The question is not whether a determined expert with unlimited resources could do it. It is whether re-identification is reasonably likely for someone motivated to try, using records and tools realistically available to them.

What do you look for in my own words?

Four kinds of bridge, each handled differently.

  1. A direct identifier written into the text, such as a name, email address, phone number, handle or account number. We remove it. “Hi, I’m [name], do these tablets cause nausea?” becomes “do these tablets cause nausea?”
  2. A combination of smaller details that together narrow to one person: age, plus job, plus town, plus family situation. We generalise it. “As a 34-year-old nurse in Brighton with two kids…” becomes “as a healthcare worker and parent…” The research point survives. The combination no longer points at anyone.
  3. A story specific enough to be a fingerprint. A detailed account of a real chain of events, with no name attached, but with enough true detail that someone who knew the situation would recognise it. We reduce it to its theme, or it does not leave. A long account of a particular workplace accident and the claim that followed becomes “a denied workplace-injury compensation claim”. A person, not software, makes this call.
  4. An exact phrase that could be matched against someone else’s records. This one is less obvious. A very specific string of text can act like a key. If the same words, at the same moment, also sit in a log another company holds against your identity, matching the two is easy. We rewrite and generalise so the exact match is broken. “Best price 2019 VW Tiguan SEL in BN1 3GH” becomes “used mid-size SUV pricing”.

(Examples are illustrative and fictionalised.)

After any change we run the checks again on the result. Something counts as scrubbed only when a fresh pass finds no bridges at all.

Is rare the same as identifying?

No. This is a common misunderstanding.

Roughly one in seven searches people type each day has never been typed before, and almost none of them identify anyone. A one-of-a-kind question about a firmware update on a pair of headphones is unique and completely anonymous.

So we do not judge your text by how unusual it is. We read it for whether it carries a bridge to a person. Rarity matters in one way only: an unusual exact phrase is easier to match against someone else’s records. We handle that as the fourth kind of bridge above.

What happens to something that can't be cleaned up?

Sometimes what makes text valuable for research is the same thing that makes it identifying. Generalising it away would leave nothing to study.

When that happens we do not force it out the door. The material stays inside our secure research environment. What can leave is aggregate findings, patterns across many people, or a trained model, never the underlying text. We would rather hold something back than release a weakened version and call it anonymous.

What about journeys?

A journey is a time-ordered sequence of your activity: what you did, in what order, and when.

Journeys need extra care, because a sequence of timestamped events is close to a fingerprint. A handful of points in time and place can be enough to pick one person out of a very large crowd. Cleaning up the words in a journey does not fix this, because the pattern itself is what identifies.

So for journeys:

  • We scrub identifiers and references as described above.
  • Journeys contain no location data.
  • We share them only with research recipients who do not hold a dataset the sequence could be matched against, so the pattern has nothing to match to. Those recipients use the detail to find broad behavioural patterns, not to look at individuals.
  • Detailed individual journeys are not made available for advertising on this basis. If your journey data is ever used for advertising, that happens only through the separate opt-in described in the Privacy Notice, and it runs on consent rather than on an anonymity argument.

What about my age, location and other profile details?

Free text and structured details need different tools, and both have to pass.

For structured fields such as age band, region or demographics, we use established statistical measures:

  • No small groups. Any combination of details we release has to be shared by at least five people, so nobody stands alone in a category of one. This is the same threshold the NHS uses for health data.
  • No uniform groups. Within any such group, sensitive attributes have to vary. A group of five people who all share one health condition would reveal that condition about every one of them, even though no individual was singled out.
  • A tested ceiling on leftover risk. We simulate attempts to single out, link and infer, and require the estimated residual risk to stay under a set limit.

A release has to clear the statistical checks and the free-text checks. Passing the numbers does not help if an identifying story survived in the text.

So is scrubbed data “anonymous”?

The answer has two parts.

To the recipient: yes. Once the bridges are gone, someone receiving the scrubbed output has no route back to you. That is what scrubbing is for, and it is what lets us share research findings safely.

To us: no. We still hold the key that links your data to your account. Everything we hold about you therefore remains your personal data, you keep every right over it, and we remain fully responsible for it. We do not use scrubbing to sidestep that.

Two further points:

  • Anonymity depends on who holds the data. The same output can be safely anonymous with one recipient and not with another, depending on what else they already hold. So we assess each release against its actual recipient rather than stamping a dataset “anonymous” permanently.
  • Anonymity is a moment in time. If new matching datasets appear, an assessment can stop holding. We re-check when the recipient, the data, or the available matching data changes.

What about health and other sensitive information?

Some sources you can share, such as search history, browsing history and AI-assistant conversations, come to us in full and sometimes contain sensitive information. Most often that is health: questions about symptoms, medications or treatments. Occasionally it touches on beliefs or other sensitive areas.

We do not ask for this, and our research never depends on any one person’s sensitive information. Where it is present:

  • We use it only for scientific and statistical research, under the specific condition UK data protection law provides for research, with the safeguards that come with it.
  • We do not use it to make decisions about you as an individual.
  • It is never used for advertising or marketing, and never made available through our data marketplace. There are no exceptions to this.
  • It leaves Measure only after scrubbing, and where scrubbing would destroy it, it does not leave at all.

Who decides, and does a person actually check?

Both software and people.

Automated tools do the bulk work: finding identifiers, flagging combinations, spotting candidates that need a closer look. They are good at volume and consistency.

The final decision on a release is a written human judgement rather than a score. Someone defines what is being released and to whom, tries to break it along all three routes above, judges whether the remaining risk is remote, writes that down and signs it. Anything involving a self-disclosing story gets human review before it can leave. That category cannot be handled safely by software alone.

We keep those assessments as our record. The law expects us to be able to demonstrate that data is anonymous, not merely to assert it.

What we don't claim

  • We do not claim scrubbing makes data anonymous to us. It does not, and it is not meant to.
  • We do not claim automated detection is perfect. That is why the protection is not detection alone: verbatim free text does not leave through high-risk channels in the first place, so a missed item cannot escape that way.
  • We do not claim an anonymity decision is permanent. It is a judgement about a specific recipient at a specific time, and we revisit it.
  • We do not use “it’s anonymised” as a reason to do things with your data you have not agreed to.

Is any of this used for advertising?

Not unless you specifically opt in.

Research use and advertising use are separate. Advertising requires its own explicit opt-in, which you can give or withdraw on its own. If you do not opt in, your data is never used for advertising. Sensitive information is excluded from advertising use whether you opt in or not.

See the “Advertising and ad-targeting use of data” section of the Privacy Notice for the detail.

What if I change my mind?

You can stop at any time. Use the Stop Sharing control, remove the Browser Extension, uninstall the app, or contact us. Stopping takes effect immediately for any further collection.

For data already collected, you can ask us to delete it. Where you withdraw from research collection, we stop processing and sever your behavioural data from the details that identify you, so it no longer points to you.

Your full rights, including access, correction, erasure, objection, restriction, portability and withdrawal of consent, are set out in the “Your legal rights” section of the Privacy Notice.

Questions

Email privacy@measureprotocol.com. If you are not satisfied with how we have handled something, you can complain to the Information Commissioner’s Office at ico.org.uk, though we would appreciate the chance to sort it out first.