A publisher decides it does not want its archive used to train AI models, and goes looking for the way to say so. It finds, instead, four ways. It can publish a robots.txt directive in the IETF’s new Content-Usage syntax. It can drop a tdmrep.json file in a well-known location, the W3C’s mechanism. It can serve an RSL license document over five different channels at once. And it can embed an IPTC Data Mining property inside every image so the signal travels with the file. So it does all four and ends the afternoon believing it has stated its preference about as loudly as anyone could.
It has not stated one preference. It has stated four overlapping, non-identical preferences, on four incompatible surfaces, with no rule anywhere for which one wins if a crawler reads two of them and they disagree — and, underneath all of it, no commitment from a single major model developer to read any of them at all. The four standards did not converge into one signal a publisher can rely on. They layered, faster than they converged. So the diligent publisher who deployed everything is not four times as protected as one who deployed nothing. It is exactly as protected: it has expressed a wish.
The watermark and provenance posts on this blog cover the output side — proving an asset was AI-made, and whether two authenticity layers agree. This is the input side: the signals you attach to your own content to declare it off-limits for training, and the composition failure that leaves you without a single enforceable one.
Four vocabularies that mean almost the same thing and share almost no tokens
Start with the part that sounds like it cannot be true: there is no agreed word for “do not train on this.” There are four, and they do not match.
The IETF’s draft has the most institutional weight behind it. A Vocabulary For Expressing AI Usage Preferences sorts AI usage into five categories — all, train-ai, train-genai, ai-use, search — where train-genai is the narrower “training general purpose AI models that have the capacity to generate text, images or other forms of synthetic content,” and ai-use is inference-time use of an asset as input to an already-trained model.
Really Simple Licensing, published by a group with editors from Conde Nast, Ziff Davis, Yahoo, Automattic, and O’Reilly, also has five usage tokens — different ones. RSL 1.0 splits AI usage into ai-all, ai-train, ai-input, ai-index, and search: ai-input is “input into AI models, including retrieval-augmented generation, grounding,” and ai-index is “inclusion in an AI system’s internal index or retrieval database.” The overlap is obvious and the mismatch is total. RSL’s ai-input and AIPREF’s ai-use gesture at the same inference-time use, but they are different strings and a parser that knows one does not know the other. RSL breaks out ai-index where AIPREF folds indexing into search; AIPREF splits train-ai from train-genai where RSL has a single ai-train. Same intent, incompatible vocabulary.
The W3C’s effort is a third one. W3C’s TDM Reservation Protocol group records, in its own April 2025 notes, that “the notion of tdm policy is specific to this effort” — a separate notion of policy again, neither RSL’s nor AIPREF’s.
The fourth is older and stranger, because it predates the current rush by two years and comes from photo metadata. IPTC Photo Metadata Standard 2023.1 adds a “Data Mining” property with nine controlled values adopted from the PLUS Coalition — DMI-PROHIBITED, DMI-PROHIBITED-AIMLTRAINING, DMI-PROHIBITED-GENAIMLTRAINING, and six more, some deferring to an external constraint or rights expression. DMI-PROHIBITED-GENAIMLTRAINING, AIPREF’s train-genai=n, and an RSL ai-train prohibition are three ways to write the same wish, and no two are the same token.
Four vocabularies, each internally reasonable, none a superset of another. A crawler that implemented one and reads content carrying another sees a preference it has no parser for — which, for a signal whose entire purpose is to be machine-read, is the same as seeing nothing.
Four attachment surfaces that do not overlap
If the vocabularies merely differed in spelling, a translation table could paper over it. They do not: the four standards also attach to content in fundamentally different places, so a publisher who serves one is, on the channel the others use, silent.
RSL alone supports five discovery methods. Per the RSL 1.0 specification, a license can be advertised through robots.txt with a License: directive, through an HTTP Link: header, through HTML link or script elements, through RSS feed elements, and through metadata embedded in the media file in EPUB, XMP, or ID3 form. AIPREF attaches through the robots layer alone: the IETF’s progress update describes preferences expressed “using the new Content-Usage mechanism in robots.txt and/or HTTP headers.” TDMRep’s primary surface is a well-known JSON file — the W3C notes describe tdmrep.json as “a well-known file on the web server where content files are stored.” And IPTC’s mechanism is different in kind from all three: the IPTC announcement states that because the Data Mining fields are embedded in the image file, “the information will be retained even after an image is moved.”
That last difference breaks any hope of a single surface covering the others. A publisher who expresses everything through robots.txt and HTTP headers — the AIPREF approach — has said nothing inside the file. The moment an image is downloaded, re-hosted, or scraped from somewhere the web server does not control, the robots-layer preference is gone, because that layer lives at the origin, not in the bytes. IPTC’s embedded property survives the move; the robots directive does not. Conversely, embed IPTC metadata and nothing else and you have expressed no preference to a crawler that reads robots.txt and never opens the pixels. The surfaces are not redundant copies of one signal. Each covers a channel the others miss — which is precisely why the authoritative advice is to deploy all of them, and precisely why no one of them is sufficient.
The standards bodies are not pretending this is solved
This is not an outsider’s complaint; the authors say it in their own documents, which is the strongest evidence the fragmentation is real rather than rhetorical. The IETF group is explicit that AIPREF exists because the layer is fragmented: its stated motivation, per the progress update, is to address “the current hodgepodge of proprietary practices” in AI content usage, under “the pressure of impending policy requirements for technical means of opt-out from AI crawlers.” That is the working group describing the proliferation of competing signals as the problem it was chartered to solve — while, by shipping a vocabulary and attachment mechanism of its own, adding a fourth.
The W3C group is equally candid, and names AIPREF directly as a competitor. Its April 2025 notes call IETF AIPREF “the new kid in town, focusing on an evolution of robots.txt,” and float the possibility that “an evolution of the latter could satisfy users of the former, and a reference to the IETF specification would then replace the tdmrep.json technique” — the W3C openly weighing whether its own mechanism should be replaced by a pointer to the IETF’s. Convergence here is a live negotiation between bodies, not a settled precedence rule a publisher can rely on today, with the same notes listing still more efforts: C2PA, Spawning AI, TDMAI.
Even within a single standard, conflicts are expected and resolved by an invented rule. AIPREF specifies one: per the vocabulary draft, “a preference that is expressed for the more general category applies if no preference is expressed for the more specific category.” But that rule governs categories within AIPREF. It says nothing about an AIPREF signal, an RSL signal, and an IPTC value landing on the same asset and disagreeing. The Search Engine Land coverage notes that even AIPREF alone needs “a standard method for reconciling multiple expressions of preferences.” If a single standard needs a reconciliation algorithm because its own expressions routinely conflict, the cross-standard case — four vocabularies, four surfaces, no shared arbiter — has no algorithm at all.
The honest answer to “how do I opt out” is twelve techniques, and it still does not bind
Ask what a publisher should actually do to reserve data-mining rights, and the most authoritative answer does not name a standard. It names a stack. IPTC’s Generative AI Opt-Out Best Practice Recommendations tell publishers how to reserve those rights “using currently available technologies,” and the list runs twelve long — spanning the IPTC Photo Metadata Standard, the Video Metadata Hub, schema.org, robots.txt, tdmrep.json, trust.txt, HTML meta tags, HTTP headers, the IPTC Data Mining property, and C2PA-signed content assertions. That is the considered recommendation from a standards body: not “use this one,” but deploy a dozen distinct mechanisms across four families, because no single one carries the preference by itself. The twelve-technique stack is the acknowledgment, in list form, that the signal cannot be expressed once.
And the recommendation concedes the punchline itself: the same guidance, per IPTC, notes that “crawlers that ignore robots.txt and other metadata” exist. Even the maximal stack produces an opt-out a crawler is free to disregard — the most complete possible expression of a preference that remains, on the other end, optional to honor.
The layer below all of this: nobody has agreed to obey
Suppose the vocabularies matched, the surfaces composed, and a clean cross-standard precedence rule existed. The publisher would still not have an enforceable preference, because the binding layer — a commitment from model developers to read and respect the signal — does not exist for any of the four. The standards say so themselves, in the plainest language in this whole post. The IETF co-chairs write, in the progress update, that “a preference on content is only the preference of the person who put it there; it is not legally enforceable on its own,” and that “preferences aren’t technically enforced. That’s something that the legal regime that they operate in will need to take up.” The vocabulary draft disclaims force at the spec level too — an entity “MAY choose to respect these preferences.” RSL is no different: the specification “contains no statement regarding current adoption by major AI companies or enforcement mechanisms,” and notes enforcement “depends on individual publisher implementation.” Each standard, in its own text, describes the opt-out it defines as a request rather than an obligation.
The field evidence matches the disclaimers. The Search Engine Land report observes that “there’s no clear evidence that AI companies follow llms.txt or honor its rules” and that “Google explicitly said it doesn’t support llms.txt” — a major developer publicly declining one machine-readable signal outright. The same piece notes that as of late 2025 “nothing from the group is final yet,” so a publisher adopting AIPREF today is betting on an unfinished standard plus future, uncommitted adoption. There is no signed commitment underneath any of the four to turn a published preference into an obligation.
Where this argument could be wrong
The skeptical reader should push on three points, because an overstated version of this thesis would not survive them.
First, “the standards do not compose” is not “the standards are useless.” A machine-readable preference, even an unenforceable one, is the substrate a future legal regime or voluntary commitment would attach to — and the IETF co-chairs frame it exactly that way, as the technical half of a problem whose other half is law. The claim here is narrow: deploying the signals today does not yield one unambiguous, enforceable preference. A preference you can point to in a dispute is worth more than none. The argument is about enforceability and composition, not value.
Second, convergence is not impossible, and the bodies are pursuing it. The W3C’s openness to replacing tdmrep.json with a reference to the IETF spec is a convergence move, and AIPREF’s stated purpose is to consolidate the “hodgepodge.” It is plausible that in a few years one robots-layer mechanism dominates and the others reference it. The thesis is about the present tense — signals deployed now, into a layer that fragmented faster than it converged — not a forecast that it stays fragmented forever.
Third, the four are not equally finished, and lumping them flattens a real difference. IPTC 2023.1 has been published since October 2023; RSL shipped a 1.0 on December 10, 2025; AIPREF, as of late 2025, had nothing final. But maturity does not fix composition: a finished IPTC property and a finished RSL token still do not share a vocabulary, still attach to different surfaces, and still bind no one.
What a publisher can honestly claim
No arrangement of the four yields one enforceable signal, because the gap is not in the publisher’s deployment — it is the absence of a shared vocabulary, a shared surface, a cross-standard precedence rule, and a binding commitment. What a publisher can do is be precise about what the work buys.
Deploy the stack anyway, matching each vocabulary to a channel. The IPTC twelve-technique list is right precisely because no single surface covers the others: the robots-layer AIPREF mechanism covers crawlers at your origin, the embedded IPTC property covers images after they leave it because it survives the move, the RSS and HTML channels cover syndication. You deploy all of them not because the union becomes enforceable, but because each closes a gap the others leave open. Pin versions, and track AIPREF as the moving, unfinished target it is; RSL buys a license-and-enforcement model the others lack but binds only conformant clients. No one tool is the whole answer.
Keep enforceability separate from expression — and say the true sentence. The signal layer is what you control. Whether the signal binds is a question of law and of voluntary commitment from model developers, and on today’s evidence neither is in place. So the accurate claim is “we have expressed, on every available channel, that we reserve data-mining rights,” not “we have opted out” — the second implies an obligation no standard creates and no major developer has accepted. Conflating the two is exactly the error the standards’ authors went out of their way, in their own specs, to warn against. It is the same discipline the output side demands: a watermark that does not survive a paraphrase is not the guarantee it reads as, and a signal no one has agreed to obey is not the opt-out it reads as.
Reading list
- RSL 1.0 Specification — defines its own five-token vocabulary (
ai-all/ai-train/ai-input/ai-index/search), five discovery surfaces, and three optional enforcement protocols, while stating no major AI company has adopted it. - A Vocabulary For Expressing AI Usage Preferences — the IETF draft vocabulary (
all/train-ai/train-genai/ai-use/search) with a within-standard precedence rule and an explicit disclaimer that compliance is optional and overridable. - Progress on AI Preferences — the AIPREF co-chairs concede a published preference “is not legally enforceable on its own” and isn’t technically enforced, and frame the whole effort as cleaning up a “hodgepodge of proprietary practices.”
- New web standards could redefine how AI models use your content — reports that nothing from the working group is final, that there is no evidence AI companies honor such signals, and that Google explicitly does not support llms.txt.
- W3C TDM Reservation Protocol Community Group Notes, April 15th 2025 — a third vocabulary with a notion of policy “specific to this effort,” delivered via a well-known JSON file, with the group openly debating whether to defer to the IETF’s “new kid in town.”
- Rights holders can exclude images from generative AI with IPTC Photo Metadata Standard 2023.1 — a fourth vocabulary of nine PLUS Data-Mining values, embedded in the file so it survives being moved, a different attachment surface in kind from the web-server standards.
- IPTC publishes best-practice guidance on Generative AI opt-out for publishers — the authoritative how-to answer is twelve distinct techniques across four families at once, with the concession that crawlers ignoring
robots.txt and metadata exist.
The publisher did everything right and ended with nothing enforceable, and that is not a failure of diligence — it is the shape of a layer that fragmented faster than it converged. Four vocabularies whose tokens do not line up, four attachment surfaces that each cover a channel the others miss, no rule for which one wins when they collide, and beneath all of it not a single binding commitment from the developers the signals are aimed at. Deploy the whole stack, because each piece closes a gap the others leave open. Then say the true sentence and not the comfortable one: you have expressed, on every channel available, that you reserve these rights — you have not opted out, because opting out is a thing the other side has to agree to, and so far the standards’ own authors are the ones telling you it hasn’t.