Rights, consent & permitted use

Whether you are allowed to use the data this way, and where it may sit.

15documents on this topic
12organizations represented
2issues named
1sourced citations
0sourced statistics

The state of it

One of 5 topics within Data readiness.

The question of whether you are allowed to use a dataset this way is asked late, if at all, and it is asked of the wrong people. Data acquired for one purpose - servicing a customer, running a transaction - is not automatically available to train on, embed, or retrieve into an answer given to somebody else.

Three failures recur. Data gathered under one consent is reused under another. Data licensed from a third party carries terms about derived works that nobody read before it went into an index. And personal data ends up embedded in a vector store, where it is genuinely difficult to delete on request, because an embedding is not a row.

That last one deserves particular care. Deletion obligations were designed against databases where a record can be removed. An embedding derived from a person's data sits in an index that has to be rebuilt to remove it, and any answer already cached from it is a separate problem again. This is worth deciding before the index is built rather than after the request arrives.

The issues, by agreement

How many independent organizations name each issue as a problem. An issue is only as real as the number of separate publishers that identify it, so the count is the ranking. Bars are organizations, not documents. Where the count reads ours, no publisher here states the issue and the analysis is our own.

Where they disagree

No contradictions recorded on this topic yet.

The issues in full

Each issue carries the organizations that name it, the numbers behind it, and the remedies proposed - with the concrete steps under each. Every citation points at a section of a named document, so any count here can be checked.

Issue 01Our analysis2026 evidence

Access to a dataset is not the same as permission to train on it

A dataset exists, the team has access, and access gets treated as permission. Whether the basis it was collected under allows this particular use is a question the sources here touch on obliquely and none of them settles.

The gap is organizational rather than legal: the people who know what the data may be used for sit in privacy, procurement or legal, and the people building the pipeline have access without ever meeting them. Third-party data is the sharpest version, because a licence often speaks specifically to derived works and a vector index is a derived work, but first-party data collected under a narrow consent raises the same question. The cost of asking is a meeting; the cost of not asking lands after the system is in production. No organization in this index names this substitution of access for permission. It is set out here as our own reading, and the consensus count is zero for that reason.

How to fix it — 1 approach, 3 steps

Establish permitted use before a dataset enters an index

No dataset is embedded, trained on or retrieved into an answer until somebody has confirmed what its collection basis or licence allows.

Done when Every dataset in a live pipeline records the basis it was collected or licensed under, anything third-party carries a written permitted-use position, and the pipeline refuses a dataset that has neither.

  1. For each dataset in a live pipeline, record the basis it was collected or licensed under.0-30 daysHead of privacy
  2. Get a written permitted-use position for anything third-party before it is indexed.30-90 daysGeneral counsel
  3. Make permitted use a gate in the pipeline pattern, not a review after the fact.90-180 daysHead of data engineering
The evidence — 1 document
OrganizationDocumentPosition
McKinsey & CompanyConsultancy · June 2026Data readiness for scaling AI impactOur reading Proposes that control travel with content through the pipeline alongside meaning and lineage, which is what allows a restriction to survive into the index.Control carried with the contentproposes a fix

Issue 02Our analysis

Whether a deletion request reaches what has already been embedded

Personal data is removed from the system of record on request. Whether the embedding derived from it goes too, and whether anyone has checked, is a question the sources here do not answer.

Deletion obligations were written against databases, where removing a record removes the data. An embedding is not a row: it is a derived representation sitting in an index that generally has to be rebuilt to remove anything, and any cached answer produced from it is a further copy. The workable answers are all architectural and all cheaper before the index exists - keeping the link from every embedding back to its source record, partitioning indexes so a rebuild is bounded, and deciding a rebuild cadence in advance. Discovering the question when the first request arrives is the expensive path. No organization in this index states that deletion fails to reach an index. The question is raised here because the architecture makes it a real one, not because a publisher reported it, and the consensus count is zero for that reason.

How to fix it — 1 approach, 3 steps

Decide how deletion reaches the index before building it

Keep every embedding traceable to its source record, bound the cost of a rebuild by partitioning, and agree the response time in advance.

Done when A deletion request has been traced through to the vector indexes and shown to reach them, each embedding links back to its source record, and the indexes are partitioned so removal rebuilds a segment rather than the whole.

  1. Establish whether a deletion request today would reach your vector indexes. Usually not.0-30 daysHead of privacy
  2. Keep a link from each embedding to the source record it derives from.30-90 daysHead of data engineering
  3. Partition indexes so removal rebuilds a segment rather than everything.90-180 daysHead of architecture
The evidence — 0 documents
OrganizationDocumentPosition

Who is represented

This dossier is drawn from 12 organizations working on the subject, 1 of which are cited directly in the issues above.

Consultancy — 2

McKinsey & Company 3 UST 1

Institution — 3

Association of Corporate Counsel 1 Cloud Security Alliance 1 NIST 1

Academic — 1

Carnegie Mellon SEI 1

Hyperscaler — 4

AWS 2 Microsoft 2 Google Cloud 1 IBM 1

Frontier lab — 1

OpenAI 1

Enterprise — 1

Grab 1