Embed

AI Text Mark – invisible Unicode watermarking for text provenance and attribution

Embed

Hide an invisible watermark in your text – a name, a label, an ID – that travels with every copy and paste

Mark your text. Question everyone else's.

AI Text Mark hides a short, invisible, machine-readable message in text you share – your name, a client ID, a label like "AI-assisted" – and reads it back after copy and paste. Tag your work, label model output, or find out which copy of a document walked away.

And for text you didn't write, the AI content analyzer measures how it reads against thousands of human- and machine-written documents – and shows you the evidence, not a made-up percentage.

Was this written by AI? Detect – the AI content analyzer Paste any text and see how it measures against real human writing: twelve statistics, each shown with how much machine text it actually catches and where it misleads. Other tools sell you a confident percentage; we show you measurements you can check – and we tell you what they're worth. That honesty is the product. Analyze a text →

The honest limits, up front: any text watermark dies when the text is retyped, and nothing on this site claims to prove who wrote something – no text tool can. The frame format is openly specified (TMK2 · TME1), the same engine is available as a REST API, the failure modes are documented next to every result, and your text is never stored.

Recommended · hybrid · interleave ×3 · ECC medium

Advanced options
Capacity –

Fill in the cover text, then press Embed watermark.

⌘ ⏎ runs it from anywhere

            

            

Optional. Token ids from your tokenizer plus a key you hold. Not Gemini, not Claude. Leave blank for a normal inspect.

Lists every invisible or substituted character a text contains, and where.

Check provenance reads TMK2, C2PA, StegCloak and SNOW, and names who can check Claude or Gemini rather than guessing.

Paste text on the left, then press Analyze or Check provenance.


            

        

Session

Not signed in

API

The short version. The full reference has every request and response.

POST /api/v1/auth/register
POST /api/v1/auth/login
GET  /api/v1/auth/me
POST /api/v1/keys
GET  /api/v1/keys
DELETE /api/v1/keys/:id
POST /api/v1/embed|extract|analyze|strip|capacity|compare
POST /api/v1/provenance
POST /api/v1/evidence · /detect · /report
GET  /api/v1/report/key
GET  /api/v1/watermarks

Provenance body (additive):
  { text, password?, stegcloakPassword?,
    askVendors?, synthid?, kgw? }
  → { read, vendor, route, hints, caveats }

read: TMK2, C2PA §A.8, StegCloak, SNOW,
      unnamed carriers, caller-key SynthID/KGW
vendor: issuer statements (never a local decode)
route: Claude · Gemini · Copilot · Meta · …
synthid/kgw need tokenIds + your key – not Gemini

Techniques: zero-width, unicode-tags, whitespace,
trailing-whitespace, homoglyph, punctuation,
variation-selector, hybrid

Frame TMK2 · ECC TME1 (default medium)

How text watermarking actually works

Two different technologies share the name “text watermark”, and almost every argument about the subject is really two people describing different ones. Here is the distinction, what removes each, what the regulation actually says – and where our own product sits, including what it cannot do.

Updated 16 August 2026 · sources linked throughout · the demos run entirely in your browser

1. There are two kinds, and they behave almost oppositely

Generation-time watermarking modifies the model's sampling. Where several continuations are roughly equally good, the sampler prefers ones that agree with a pseudorandom key. The signal lives in which words were chosen. This is what Google's SynthID-Text does in Gemini – and, as of Anthropic's July 2026 announcement, what Claude uses too: the same SynthID-Text technique, licensed from DeepMind.

Post-hoc watermarking is applied after the text exists. Format-based schemes – the class this product belongs to – hide a payload in the encoding rather than the words: invisible Unicode characters, substituted spaces, lookalike letters. The words are untouched.

Generation-timePost-hoc format-based
Signal lives inWord choiceCharacter encoding
Needs decoder controlYesNo – works on any text
Survives retyping / OCRYesNo, never
Survives a regex stripYesNo
Survives paraphraseDegradesNo
Carries a payloadUsually zero-bitYes – arbitrary bytes
Who can detectKey holder onlyAnyone, if specified

Neither is better. They answer different questions. Generation-time answers “did a model write this?”; post-hoc answers “which copy of this document is this, and who received it?”

Try it – the thumb on the scale

A model picks the next word from several acceptable candidates. In the green/red-list construction, a secret key sorts each position's candidates into two sets – depending on the previous word, so the same word flips colour with context – and the sampler quietly leans green. One choice proves nothing; the lean only shows up in the count.

…the harbour lights ____

no choices yet

It is not a word list – the same word flips colour with context. “trembled” after…

Try it – detection is counting, and only the key holder can count

This paragraph was written with the issuer's key leaning on the scale. Switch keys and the same words stop meaning anything – that is why every deployed generation-time watermark is detectable only by its issuer. Then edit the paragraph freely and watch the count fall toward chance: you are doing, by hand, exactly what a paraphrasing tool does.

Try it – a post-hoc mark is really in there

The two sentences below are identical to the eye. The first was marked by this site's engine – zero-width characters carrying the payload id=42 – and the second was retyped by hand. Reveal the carriers, or copy the marked sentence and read the payload back in Extract. The honest limit is visible too: the retyped copy carries nothing, because retyping is all it ever takes.

Marked by the engine

Retyped by hand

Try it – SynthID's real algorithm, ported to this page

Everything above is our miniature. This is the real thing – and since Anthropic adopted SynthID-Text for Claude, it is the mechanism behind two of the three frontier labs' watermarks: a JavaScript port of DeepMind's SynthID-Text reference implementation – the same 64-bit hash, the same g-value derivation, the same tournament sampling – verified bit-for-bit against their Python before it shipped. Where our demo colours a word with one bit, SynthID hashes the previous four tokens through 30 independent keys, so each candidate carries thirty green/red votes at once, and the sampler runs a tournament between candidates rather than a simple nudge.

context

Two consequences fall straight out of their code. Detection needs no model – you re-derive the g-values from the text and take a mean: unwatermarked text averages 0.5, and their own arithmetic caps watermarked text at 0.5 + 0.25·(1 − 1/V), so even a perfect signal moves the mean by a quarter. And their processor skips any four-token context it has already watermarked, so repetitive or short text carries less signal by design. Detection is statistics either way – which is why the honest output is a confidence, never a certificate.

Try it – read the frame like a machine would

“Openly specified” is not a slogan; it is the 17 bytes below. The 66 invisible operators in the sentence above decode – base-5 digits to bits to bytes – into one TMK2 frame, and because the format is published, anyone can write this decoder. Every deployed generation-time watermark is a black box its issuer alone can read; this one is legible by design.

2. “Undetectable” is a precise claim, not marketing

A common objection is that watermarking must degrade output. For generation-time schemes this is answerable formally. Christ, Gunn and Zamir (2023) define a watermark as undetectable if no polynomial-time adversary without the key can distinguish watermarked from unwatermarked output – and construct one from one-way functions alone. Under that definition the quality question is settled by construction: you cannot tell, so there is nothing to degrade.

Post-hoc Unicode marking makes the opposite trade. It never changes a word, so meaning is untouched – but it is trivially detectable. Frontier models identify these schemes on sight, and a tweet-length input can grow several times in length once invisible characters are added.

3. Yes, paraphrasing removes it

This is true of both classes and nobody serious disputes it. The interesting part is cost.

As Jonas Geiping puts it, a light edit is not enough for a generation-time watermark keyed on overlapping n-grams: none of the original n-grams the detector keys on may survive, which for a long document means rewriting nearly all of it. Kirchenbauer et al. found watermarks still detectable after strong human paraphrasing given roughly 800 tokens.

For post-hoc Unicode marking the cost is close to zero: one regular expression, or retyping the text. We publish that rather than hide it.

What removes an invisible Unicode mark
  • A regex over default-ignorable code points – takes out zero-width, Unicode Tags and variation selectors together, because they are one class, not three.
  • NFKC_Casefold, the normalisation behind identifier and loose matching, deletes the whole carrier set by definition. So does step 2 of the UTS #39 confusable check.
  • Retyping, OCR, screenshots, print-and-scan. Always, with no configuration that helps.

The "watermark removers" do not remove watermarks

A repository called watermarks-remover gathered more than ten thousand GitHub stars in its first four days by promising to strip AI provenance marks from Claude, Gemini/SynthID and OpenAI output. It does not do that. It cannot do that, and neither can any of the repositories copying it. We cloned it, read the code, and ran it against real marks. Here is what is actually inside.

The part that runs is a one-line regex. Its deterministic layer deletes invisible Unicode – zero-width characters, Tags-block runs, variation selectors, exotic spaces – and normalises the rest. That is the whole mechanism. It works on marks like ours because we say on this very page that one regex over the ignorable class removes them; it is the "strip all invisibles" button in the demo below, restated in Python. It is not research, it is not new, and it has nothing to do with any frontier lab's watermark – the Claude and Gemini marks are not carried by invisible characters, and OpenAI ships no text watermark to strip. Run it on watermarked Claude output and the watermark is untouched, because it was never in the characters this tool deletes.

Against the marks it advertises, it has no removal at all. A generation-time watermark – SynthID-Text, and now Claude's – lives in which words the model chose. There is no character to strip and no filter to apply. So for those, the repository ships a prompt: it asks some other language model to paraphrase the text ("use substantially different wording at the token level… replace both content words and function words"), or translate it out and back, or reduce it to bullets and write it again. That is not a watermark remover. That is a paraphraser with a watermark remover's name – and it does exactly what the counting demo above lets you do by hand: replace enough words and there is nothing left to count. Its own removal-matrix.md, in a column headed "Verifiable today?", answers for the statistical marks: "No without vendor key/detector." Read that plainly. The tool cannot see the watermark it claims to remove, so it cannot tell you whether it removed it. And a rewrite by another model routinely replaces one vendor's watermark with another vendor's – the repository's own cross-vendor table exists because its authors know that.

The rest of it is metadata deletion and an invented score. The C2PA half strips Content Credentials from images and documents – which works, because C2PA metadata is easy to delete and its own designers say so; the imperceptible watermark that C2PA soft-binding falls back to is untouched. The bundled "stylometry" scorer is a hand-weighted composite – 45% sentence burstiness, 45% a list of AI-ish phrases, 10% vocabulary richness – producing one number and an exit code that says "suspicious". That is precisely the invented score our Detect tab refuses to print, and the reason it refuses is measured on the same tab: those statistics overlap heavily between human and machine text.

The repository is competently built, and its own reference notes are far more honest than its title – every caveat above is in there. None of it made the star count, and none of it survives being repeated on social media as "this removes the Claude watermark". It does not. Nothing does, short of what Anthropic itself describes: a complete rewrite where every word is replaced.

The plain summary of every "remover"
  • Claude, Gemini/SynthID, OpenAI: not removed. Not detectable by these tools, not strippable, and a paraphrase by another model is not removal – it is a rewrite the tool cannot verify and that may re-mark the text.
  • Post-hoc Unicode marks (this product's class): removed – by a regex we publish ourselves. Any text editor's find-and-replace does the same.
  • C2PA metadata: deleted; the soft-bound watermark it points to is untouched.
  • The "AI score": three statistics with made-up weights. Not a detector.

Try it – attack a real mark and watch what survives

The paragraph below carries our hybrid mark – the same frame written redundantly into zero-width, Unicode-tags and variation-selector channels. Every button applies the real operation in your browser (the normalisation button literally calls normalize('NFKC')), then asks the real extractor what is left. Notice the middle of the story: killing one channel is not enough, because the frame rides three of them – and then one regex takes all three at once, exactly as the callout above admits.

Pick an attack – the extractor will answer.

Try it – damage it and let the error correction fight back

Under the frame sits Reed–Solomon error correction – the arithmetic that lets a scuffed CD still play. The sentence below is marked with heavy ECC. Corrupt a few carriers at random and the payload usually survives; corrupt more and it dies. And sometimes a small hit lands on a load-bearing byte and kills it early – position matters, which is exactly how error correction behaves outside of brochures.

4. No third party can tell you which model wrote something

Including us. Every deployed generation-time text watermark is symmetric-key – detection needs the issuer's secret. Google's SynthID-Text is gated, and zero-bit: it reports watermarked or not, carrying no payload to read. Anthropic announced in July 2026 that future Claude models carry a SynthID-Text watermark and that a detection API is coming – key-gated like Google's, so it will tell you whether Claude was involved and nothing about any other model. OpenAI ships no text watermark at all.

What Anthropic's own announcement concedes

The Claude post is unusually candid about limits, and they match everything on this page. In their words: a watermark "can only determine that Claude was likely involved with the content at some point" – it "cannot distinguish 'Claude wrote this' from 'Claude heavily edited this'", it "doesn't confirm whether the text was human-written", and it "can't tell whether the text was written by a different AI". It only lands where the model has "an arbitrary choice between particular words", so it runs sparse on factual passages and code, thin on short text, and absent on human prose Claude merely proofread. On editing: "light editing probably won't remove the watermark completely; a complete rewrite where every word is replaced will" – the counting demo above, in one sentence. And by design it cannot trace text to a user, an organisation, or a conversation. That is the frontier lab with the most to gain from overclaiming, declining to.

So a tool advertising “paste text, see which AI wrote it” is guessing, and a false positive there gets a real person accused of cheating or fabricating. Our provenance check reads what is publicly specified (TMK2, C2PA, StegCloak, classic SNOW, unnamed Unicode carriers), will relay an issuer's own detector when one is callable, and will run SynthID-Text or KGW only against a key you supply. For Gemini and Claude it names who can check rather than inventing an answer.

5. What the EU AI Act actually says

Article 50(2) requires providers of generative AI systems to mark outputs in a machine-readable format and make them detectable, with solutions “effective, interoperable, robust and reliable as far as this is technically feasible”. Three details get lost:

  • It binds the provider of the system, not a tool vendor. You can buy a marking solution, but the obligation and the burden of demonstrating compliance stay with you.
  • Marking alone is not compliance. Detection has to be available too.
  • There is no certification scheme. No notified bodies, no CE route for Article 50. Any “AI Act certified” badge is decorative.

The Code of Practice published in June 2026 asks for at least two layers of machine-readable marking, on the explicit basis that no single technique currently suffices – and lists paraphrasing, homoglyphs, character insertion and deletion, and print-and-scan among the transformations a mark should survive.

Where that leaves this product

Two of those named transformations are carriers we implement, and none of our channels survives print-and-scan. We therefore do not claim AI Text Mark satisfies Article 50(2), and we would treat any vendor who does with suspicion. What it is good at is attribution and leak tracing – per-recipient fingerprinting, where the adversary is not attacking the mark – and serving as an openly specified, publicly readable marking layer, which is the one thing almost nothing else in this space is.

6. Questions that keep coming up

If a model is trained on watermarked text, does it start producing watermarked text?

No. The mark is not a phrase or a habit that could be picked up by imitation. It is a faint preference spread across thousands of individual word choices, and which choices count is decided by a secret key. There is no consistent pattern in the text for a second model to learn, so watermarked training data does not make a watermarked student.

Could a model accidentally reproduce someone else’s watermark?

No, for the same reason. Reproducing it would mean guessing the key from the output, and the number of possible keys is astronomically large. Writing in a style that looks like Claude is not the same as carrying Claude’s watermark.

Does watermarking make a model duller or more repetitive?

Not in the way people expect. The sampler only intervenes where several continuations are about equally good, and it nudges rather than forces. If anything it makes successive generations slightly more varied, because it breaks ties that would otherwise fall the same way each time.

Can one AI agent spot another one by its watermark?

Not by reading the text. Without the key the mark is indistinguishable from ordinary randomness in how the model picked its words. Given access to a detector, yes – but that is a matter of who holds the key, not of what the text looks like.

If text carries a watermark, does that prove who wrote it?

Only when the mark is authenticated. An unauthenticated mark can be forged by anyone who understands the format, and it can be lifted out of one document and pasted into another that the signer never saw. That is why extraction here reports an authenticated flag, and why supplying a password makes it refuse unauthenticated marks unless you ask for them explicitly. Treat an unauthenticated mark as a label someone attached, not as evidence.

Will one of the GitHub "watermark remover" repositories strip a Claude or SynthID mark?

No. Not Claude's, not Gemini's, not OpenAI's – see section 3. The only thing those tools remove is invisible Unicode, which frontier-lab watermarks do not use; against a generation-time mark they hand your text to another model to paraphrase, which is not removal, cannot be verified without the vendor's key, and may re-mark the text.

One of these two words is not what it looks like – which?

A homoglyph mark swaps letters for lookalikes from other alphabets. It survives the invisible-character strippers above – and is caught instantly by a mixed-script check, which is why it is a tracing tool, not a hiding place.

Both render with your fonts, right now, in this sentence's typeface.

So is any of this worth doing?

It depends what you want from it. As a way to stop a determined person passing off AI text as their own, no – a rewrite defeats every scheme here. As a way to know which copy of a document leaked, or to give honest publishers a machine-readable signal they can attach and others can check, yes. Most of the disappointment in this field comes from expecting the first and being sold it as the second.

Frame specification

“Openly specified” means you can read the mark without us. This page is the whole format: the TMK2 payload frame, the TME1 error-correction envelope, how bytes become carrier symbols, the seven carrier alphabets, and what a conforming decoder must check. Nothing here depends on our code.

TMK2 frame version 2 · TME1 envelope version 1 · report format aitextmark-detection-report/1 · updated 3 September 2026

1. The layers

A watermark is built in three steps and read back in reverse.

payload bytes
  → TMK2 frame        magic, flags, optional salt/IV, length, body, CRC32, optional HMAC
  → TME1 envelope     optional Reed–Solomon outer code around the whole frame
  → base-N record     a length header and the bytes repacked as digits of the carrier's alphabet
  → carrier symbols   one invisible or look-alike character per digit, placed in the cover

The frame is the unit of meaning. The envelope only repairs byte errors. The record only turns bytes into digits. The carrier only decides which characters the digits become and where they go. Every carrier writes the same frame, which is why hybrid mode can lay identical copies into several channels and why any one surviving channel recovers the payload.

2. The TMK2 frame

MAGIC(4) VERSION(1) FLAGS(1) [SALT(16)] [IV(12)] LEN(2) BODY(LEN) CRC32(4) [HMAC(32)]
FieldSizeValue
MAGIC4ASCII TMK2 – 54 4D 4B 32
VERSION10x02. Any other value is rejected.
FLAGS1bit 0 COMPRESS, bit 1 ENCRYPT, bit 2 HMAC. Any other bit set → reject the frame. ENCRYPT and HMAC are never both set.
SALT16Present when either ENCRYPT or HMAC is set. Random per frame; input to the key derivation.
IV12Present only when ENCRYPT is set. Random per frame; the AES-GCM nonce.
LEN2Big-endian length of BODY in bytes, at most 65535.
BODYLENThe payload, after optional compression and optional encryption (below).
CRC324Big-endian CRC-32 of BODY only. ISO 3309 / PNG polynomial (reflected 0xEDB88320, initial 0xFFFFFFFF, final XOR 0xFFFFFFFF).
HMAC32Present only when HMAC is set. HMAC-SHA256 over every preceding byte, MAGIC through CRC32 inclusive.

Body transforms

  • Compression. The encoder tries raw DEFLATE (no zlib or gzip header, level 9) and keeps the result only if it is strictly shorter than the input, setting COMPRESS. Short identifiers therefore never grow. A decoder must refuse to inflate past 1 MiB (1,048,576 bytes); a payload is an identifier, not a document.
  • Key derivation. When a password is set, scrypt(password, SALT, N = 16384, r = 8, p = 1) produces 64 bytes. The first 32 are the AES key, the last 32 the MAC key. The password is UTF-8.
  • Encryption. With ENCRYPT, BODY is AES-256-GCM ciphertext of the (possibly compressed) payload, followed by the 16-byte GCM tag. There is no associated data. The tag is the frame's authenticator.
  • Authentication without encryption. With HMAC, the body stays in the clear and the trailing HMAC-SHA256 under the MAC key is the authenticator.
  • Neither. Without a password there is no authenticator at all. CRC32 detects accidental corruption and nothing else: anyone who has read this page can forge such a frame.

What a decoder must do

  1. Check MAGIC and VERSION; reject unknown FLAGS bits.
  2. Read the optional fields implied by the flags, then LEN, and require the buffer to hold BODY, CRC32 and, if flagged, HMAC.
  3. Verify CRC32 over BODY. A mismatch is corruption.
  4. If HMAC is set and a password is available, derive the keys and verify the tag in constant time. If ENCRYPT is set, derive the keys and decrypt; GCM verifies the tag.
  5. If COMPRESS is set, inflate with the 1 MiB ceiling.
  6. Report authenticated = true only when a GCM tag or HMAC verified. A frame that merely parsed with a good CRC is unauthenticated data, whatever it says.

Our extractor fails closed: supplying a password means an unauthenticated frame is not returned unless the caller explicitly asks for it, and when several decodable frames disagree the result is flagged as a conflict rather than silently picking one.

Worked example

The payload id=42 with no password. Five bytes do not compress, so the flags byte is zero and the frame is the 17 bytes below. The last four are the CRC-32 of id=42.

54 4D 4B 32   MAGIC "TMK2"
02            VERSION
00            FLAGS (nothing set)
00 05         LEN = 5
69 64 3D 34 32   BODY "id=42"
B9 0A 56 C1   CRC32("id=42")

As a base-5 record for the zero-width carrier: 7 length digits (17 → 0000032), then
136 bits as three 19-digit groups of 44 bits and a 2-digit tail, 66 digits in all.
Each digit d becomes the character U+2060 + d.
0000032 1224403420203031333 0000000141220334041 0432232213304420410 01

The how-it-works page decodes exactly these 66 characters live in your browser. With the default ecc: medium the same frame is first wrapped in a 41-byte TME1 envelope (8-byte header, 17 data bytes, 16 parity bytes) and the record grows accordingly.

3. The TME1 envelope

By default the finished frame is wrapped in an outer Reed–Solomon code before it reaches a carrier, so a damaged run of characters can be repaired before the CRC or authenticator is checked.

MAGIC(4) VERSION(1) NSYM(1) DATALEN(2) BLOCK … BLOCK
FieldSizeValue
MAGIC4ASCII TME1
VERSION10x01
NSYM1Parity bytes per block. Even, 2 to 64. light = 8, medium = 16 (default), heavy = 32.
DATALEN2Big-endian length of the enclosed TMK2 frame.
BLOCK≤ 255Up to 255 − NSYM frame bytes followed by NSYM parity bytes. The frame is split into consecutive blocks; the last block is shorter.

The code is Reed–Solomon over GF(28) with the QR-code field polynomial 0x11D, codeword layout message ‖ parity, as in the ZXing implementation. Each block corrects up to ⌊NSYM / 2⌋ byte errors: 8 at the default level. Damage beyond that loses the block, which is what redundant copies are for.

A decoder that finds TME1 unwraps and corrects, then parses the enclosed frame as in section 2. A decoder that finds TMK2 directly parses it with no correction; ecc: off produces that form.

4. From bytes to digits

Every carrier exposes an alphabet of N symbols, so the envelope (or bare frame) is written as a record of base-N digits:

  1. A length header: the byte length, big-endian, in the fixed number of digits needed to represent 65535 in base N (for base 5 that is 7 digits, for base 127 it is 3, for base 2 it is 16).
  2. The bytes as a bit string, packed into groups of g digits carrying b bits, where b = ⌊g · log2 N⌋ and g is chosen (up to 40, with Ng ≤ 248) to maximise b/g. Each group is the b-bit value written big-endian in g digits. For a power-of-two base this is plain bit packing; for base 5 it is 19 digits per 44 bits, 2.316 bits per symbol against the 2.322-bit ceiling.
  3. A tail group for the leftover bits, in ⌈bits / log2 N⌉ digits.

Redundant copies are further records laid end to end, or spread across the cover in interleave mode. A decoder walks records from the start; if a record fails, it scans forward for a run of digits that decodes to the frame or envelope magic and resynchronises there, so a copy after a damaged region still reads.

5. The carriers

Each carrier maps digits to characters and decides where they go. The alphabets below are exact; the rules in the last column are the ones a conforming encoder must keep so that it never rewrites the meaning of someone's text.

CarrierBaseAlphabet (digit 0, 1, …)Placement and rules
zero-width5U+2060 WORD JOINER, U+2061, U+2062, U+2063, U+2064 (the invisible operators)Inserted as a run after the first Latin letter (append-first) or as several copies at anchors spread through the cover (interleave). U+200B, U+200C and U+200D are never used: the joiners are load-bearing in Persian, Arabic, Indic scripts and every emoji sequence, and U+200B creates a line-break opportunity.
unicode-tags127U+E0001 … U+E007F in order. U+E0000 is skipped.Inserted after Latin carrier characters only, so joining-script words are never split. Default-ignorable; deleted outright by NFKC_Casefold. A leading U+FEFF in the cover is ignored by the reader.
whitespace (safe)2U+0020 SPACE, U+00A0 NO-BREAK SPACESubstitutes single inter-word spaces in place, one digit per space, left to right. Indentation, trailing spaces and space runs are skipped, so code blocks and Markdown hard breaks survive and length is unchanged.
whitespace (dense)8U+0020, U+00A0, U+202F, U+2009, U+200A, U+205F, U+2006, U+2005Same placement, three bits per space. Measurably visible next to unmarked text in proportional fonts; exists for capacity, not deniability.
trailing-whitespace2U+0020 SPACE, U+0009 TABAppended at the end of each line, at most 64 symbols per line, the classic SNOW layout. Destroyed by trim-on-save.
homoglyph2Digit 1 swaps a Latin letter for its Cyrillic look-alike: a→U+0430, c→U+0441, e→U+0435, i→U+0456, j→U+0458, o→U+043E, p→U+0440, s→U+0455, x→U+0445, y→U+0443, and capitals A→U+0410, B→U+0412, C→U+0421, E→U+0415, H→U+041D, K→U+041A, M→U+041C, O→U+041E, P→U+0420, T→U+0422, X→U+0425One bit per eligible letter, only inside word tokens that are unambiguously Latin; never in URLs, emails or paths, where a look-alike is a spoofing string. With a password the eligible positions are permuted by an HMAC of the password. The reader folds Cyrillic back only where the surrounding token is Latin, so genuine Cyrillic prose is left alone.
punctuation2Digit 1 swaps -→U+2010 HYPHEN, .→U+2024 ONE DOT LEADER, '→U+02B9, "→U+02BAOne bit per period, hyphen, apostrophe or double quote; URLs, emails, paths and filenames are skipped. The curly quotes and dashes a word processor produces (U+2018, U+2019, U+201C, U+201D, U+2013, U+2014) are neither emitted nor folded, so a document's own typography is not mistaken for a mark.
variation-selector16U+FE00 … U+FE0F (VS1 … VS16)One selector appended after an ASCII letter or digit, four bits per carrier. Emoji, CJK and any character with a defined variation sequence are skipped, because a selector there changes how it renders.

Two rules apply to every insertion carrier: characters land on grapheme-cluster boundaries, never inside a surrogate pair or between a base character and its combining marks; and a channel is written only if the whole record fits, so a frame is never truncated mid-way.

Hybrid writes the same envelope into zero-width, unicode-tags, whitespace and variation-selector by default. The first, second and fourth are all default-ignorable code points and one regular expression over that class removes them together; whitespace needs roughly 200 words of cover before a frame fits. The embed result reports survivesIgnorableStrip so a caller can see whether any non-ignorable channel was actually written.

6. Capacity accounting

Capacity is reported in the unit a caller controls: UTF-8 bytes of the string passed to embed, before compression, with everything else already subtracted.

  • carrierSlots – positions the cover offers this carrier, or null for insertion carriers that are bounded by the 65,535-byte frame rather than by the cover.
  • bitsPerSlot – the packing rate from section 4 for this alphabet.
  • maxFrameBytes – the largest serialised envelope (or bare frame) the slots can hold, length header included.
  • frameOverheadBytes – framing plus salt, IV and authenticator for the options in force: 12 bytes bare, 60 with HMAC, 56 with encryption, plus the 8-byte envelope header and parity when ECC is on.
  • maxWatermarkBytes – the guarantee. Embedding this many bytes fits; one incompressible byte more is refused.
  • bestEffortWatermarkBytes – for hybrid, what still embeds somewhere when the weakest channel is dropped.

7. Signed detection reports

POST /api/v1/report returns a JSON object with two members, report and signature. The signature is Ed25519 over the UTF-8 bytes of the canonical form of report:

  • object keys sorted by code point at every level;
  • members whose value is undefined dropped;
  • no whitespace anywhere; strings, numbers, booleans and null as JSON.stringify emits them.

The public key is served as base64 SPKI DER at GET /api/v1/report/key. signature.keySource is configured when the deployment holds a stable key, and ephemeral-process-key when it does not, in which case the report verifies only against the key that was current when it was issued. The report body carries version, issuedAt, the detector identity, a subject block with the SHA-256, byte and character counts of the submitted text, the result, and caveats derived from the result. The text itself is not in the report and is not retained.

import { createPublicKey, verify } from 'node:crypto';

const canonical = (v) =>
  v === null || typeof v !== 'object' ? (JSON.stringify(v) ?? 'null')
  : Array.isArray(v) ? `[${v.map(canonical).join(',')}]`
  : `{${Object.entries(v).filter(([, x]) => x !== undefined).sort(([a], [b]) => (a < b ? -1 : a > b ? 1 : 0))
      .map(([k, x]) => `${JSON.stringify(k)}:${canonical(x)}`).join(',')}}`;

const key = createPublicKey({ key: Buffer.from(signature.publicKey, 'base64'), format: 'der', type: 'spki' });
const ok = verify(null, Buffer.from(canonical(report), 'utf8'), key, Buffer.from(signature.value, 'base64'));

8. Known limitations and open issues

These are properties of version 2 as shipped. They are listed so nobody has to discover them; the ones marked v3 need an incompatible frame change and will arrive under a new version byte, never by silently altering the meaning of an existing field.

  • Frames are not bound to their cover (v3). An authenticated frame can be lifted out of one document and placed into another, and it will still authenticate. A signature over the payload says a holder of the password produced that payload; it says nothing about the text around it.
  • AES-GCM is used without associated data (v3). The header, including the COMPRESS flag, is not covered by the GCM tag, so an attacker can flip that bit and change how a decoder interprets an otherwise authentic body.
  • The TME1 header sits outside the Reed–Solomon protection (v3). Damage to those eight bytes is not repairable.
  • Homoglyph positions are HMAC-keyed, not KDF-keyed. Given the text, the position permutation is brute-forceable. The payload itself is separately encrypted and is not.
  • Every carrier is removable by cheap, well-known operations. A regular expression over default-ignorable code points removes zero-width, tags and variation selectors together; NFKC_Casefold deletes the entire carrier set; retyping, OCR, screenshots, print-and-scan and paraphrase destroy every channel. This format suits attribution and leak tracing, where nobody is attacking the mark. It is not a robust watermark and does not make anyone compliant with anything.

9. Reading it with the API

You do not have to write a decoder to use the format. POST /api/v1/extract reads a frame, POST /api/v1/analyze lists which carriers a text contains, and POST /api/v1/provenance also reads C2PA text manifests, StegCloak and SNOW. The API reference has every field.

API reference

Everything the lab does is one HTTP call away, on the same engine. JSON in, JSON out, no SDK required. Watermark, inspection and detection endpoints work without an account; an API key adds history and usage counters and lifts the anonymous rate limit.

Base URL https://aitextmark.com/api/v1 · responses are additive: fields are added, never renamed or removed · updated 3 September 2026

1. Conventions

  • Requests are POST with Content-Type: application/json and a JSON object body, unless marked GET or DELETE. Text goes in the body, never in the URL.
  • Errors are { "error": "message" } with a meaningful status: 400 bad input, 401 missing or invalid credentials, 403 wrong credential kind, 404, 413 too large, 429 rate-limited, 503 accounts not configured.
  • Authentication is a bearer header: Authorization: Bearer tm_… with an API key, or a session token from /auth/login. Endpoints marked optional accept either or none; with a credential, embeds are recorded to your history and every call is counted in /stats.
  • Limits. Bodies up to 4 MB. Any single text field up to 500,000 characters (/evidence and /detect: 200,000); over is 413. Anonymous callers are limited per client address, by default 120 requests a minute per server instance; over is 429 with a Retry-After header. Requests carrying a valid API key or session are never rate-limited.
  • CORS is open: Access-Control-Allow-Origin: *, preflight answered, Authorization and Content-Type allowed. A browser client on any origin can call the API with its own key.
  • Compatibility. Response shapes only grow. New fields may appear; existing ones keep their names and types. Unversioned paths (/api/embed and friends) remain as aliases of /api/v1/….
  • Privacy. Cover text, payloads and detection inputs are processed and discarded. History rows hold a 120-character cover preview and an 80-character payload preview. Usage counters hold endpoint, day, counts and latency – no bodies, no addresses.
curl -s https://aitextmark.com/api/v1/embed \
  -H 'Content-Type: application/json' \
  -H 'Authorization: Bearer tm_…' \
  -d '{"cover":"Ship the quarterly report before Friday.","watermark":"owner=acme;doc=Q2","technique":"hybrid","password":"s3cret"}'

2. Endpoints

MethodPathAuthPurpose
POST/embedoptionalWrite a watermark into cover text
POST/extractoptionalRead a TMK2 watermark back
POST/analyzeoptionalList the carriers and hidden characters a text contains
POST/stripoptionalRemove carriers
POST/capacityoptionalHow many payload bytes a cover can carry
POST/compareoptionalVisual fold of original against watermarked
POST/provenanceoptionalCross-scheme inspection: TMK2, C2PA, StegCloak, SNOW, vendor routes
POST/evidenceoptionalWriting statistics against human and machine reference corpora
POST/detectoptionalCalibrated classifier with abstention
POST/reportoptionalEd25519-signed TMK2 detection report
GET/report/key–Public key to verify reports
GET/status–Whether accounts are configured
POST/auth/register · /auth/login–Accounts
POST · GET/auth/logout · /auth/mesession · anySession
POST · GET · DELETE/keys · /keys/:idsessionAPI keys
GET/watermarks · /watermarks/:idanyEmbed history previews
GET/statsanyYour call counts, last 30 days

3. Watermarking

POST/embedauth optional
FieldTypeMeaning
coverstring, requiredThe visible text that will carry the mark.
watermarkstring, requiredThe payload. Measured in UTF-8 bytes against capacity.
techniquestringzero-width (default), unicode-tags, whitespace, trailing-whitespace, homoglyph, punctuation, variation-selector, hybrid.
passwordstringAuthenticates the frame. With a password and no encrypt, the body is encrypted by default.
encryptbooleanAES-256-GCM. Requires password. false with a password gives HMAC only.
compressbooleanDefault true; kept only when it shrinks the payload.
modestringappend-first (default) or interleave (redundant copies spread across the cover).
repeatinteger ≥ 1Redundant copies where the carrier supports them.
whitespaceProfilestringsafe (default) or dense.
hybridTechniquesstring[]Which carriers hybrid layers. Default zero-width, unicode-tags, whitespace, variation-selector.
eccstring · boolean · integermedium (default), light, heavy, off, or an even parity count 2–64.

Response. text (the watermarked cover), technique, bytesEmbedded, frameBytes, channels (the carriers that actually accepted the frame), survivesIgnorableStrip (at least one non-default-ignorable channel was written), warnings (non-fatal problems in words), capacity (as /capacity), visual (as /compare), analysis (as /analyze), and historyId (the saved row's id, or null when anonymous). A payload that does not fit is a 400 naming the shortfall.

POST/extractauth optional

text (required), technique (auto by default, or one carrier), password, whitespaceProfile.

Response. watermark (the payload as a string), technique (the carrier it was read from), valid (the frame parsed and its CRC checked), authenticated (a GCM tag or HMAC verified under the password), compressed, encrypted, bytes. Supplying a password means an unauthenticated frame is not returned, so a genuine mark can never be outranked by a forged one planted in another channel. Nothing decodable is a 400. Treat authenticated: false as data anyone could have written.

POST/analyzeauth optional

text (required).

Response. techniquesDetected (carriers present; hybrid is appended when two or more are), suspiciousChars (each with codePoint, hex, name, count, role), estimatedPayloadBytes, and counts: invisibleCharCount, spaceSubstitutions, homoglyphCount, variationSelectorCount, unicodeTagCount, trailingWhitespaceCount, punctuationCount, printableLength, rawLength. Presence of a carrier is not identification of who wrote it.

POST/stripauth optional

text (required), aggressive (boolean, default false). Returns { text }. The default strip preserves U+200C and U+200D where they are part of a script or an emoji sequence, and leaves presentation selectors on emoji alone; aggressive removes them too and will change how some text renders.

POST/capacityauth optional

cover (required), technique, password, encrypt, whitespaceProfile, ecc – the options you intend to embed with, because overhead depends on them.

Response. technique, carrierSlots (null when the channel is bounded by the frame limit rather than the cover), bitsPerSlot, maxFrameBytes, frameOverheadBytes, maxWatermarkBytes (the guarantee: this many UTF-8 bytes of payload fit), bestEffortWatermarkBytes (hybrid: what still embeds when the weakest channel is dropped), notes.

POST/compareauth optional

original (required), watermarked. Returns identicalVisually, printableOriginal, printableWatermarked, printableEqual, invisibleDelta (added or substituted code points), and codepointDiffs as { index, a, b }.

4. Provenance and detection

These three endpoints answer different questions and their outputs are never merged. /provenance asks what marks does this text carry, and who could check the ones we cannot? /evidence and /detect ask how does this text read against human and machine writing? Neither is proof of authorship; every response says so in its own caveats, and the terms of use forbid using them to accuse an individual.

POST/provenanceauth optional

text (required), password (TMK2 and StegCloak), stegcloakPassword (if different), askVendors (boolean, default false: when true a configured issuer endpoint may be called), synthid ({ tokenIds, keys, ngramLen?, hashIv?, zThreshold? }), kgw ({ tokenIds, seed, gamma?, ngram?, threshold? }). Send synthid or kgw only with token ids from your own tokenizer and a key you hold; without keys or seed the row is unreadable, never a guess.

Response. Four tiers, always all present:

  • read – marks decoded here without any vendor key. Each has scheme, status (found · absent · unreadable), detail and optional data. Schemes: AI Text Mark (TMK2), C2PA text manifest (2.4 §A.8) (parsed; the COSE signature is not verified), StegCloak, SNOW (classic), Unidentified Unicode carrier, and SynthID-Text (caller key) / KGW / Maryland (caller key) only when you supplied one.
  • vendor – issuer statements relayed, never decoded locally. status is watermarked, not-watermarked, uncertain, unavailable or not-requested.
  • route – schemes we cannot read and who can, with support of coming-soon, vendor-only, caller-key or none-shipped.
  • hints – formatting correlations such as a stray U+202F, as { signal, count, detail }. Not evidence.
  • caveats – strings, derived from the result.
POST/evidenceauth optional · 200,000 characters

text (required).

Response. wordCount, sentenceCount, paragraphCount, spansApplyToInput (whether span offsets index your text as sent), reference (how the human and machine distributions were built: generated, humanCorpus.n and sources, aiCorpus.n and note, caveat), notes, and signals: one row per statistic with name, label, what (one sentence), value, unit, direction, minWords, available, reason when it abstained, failureMode (when this misleads), position (toward-machine · toward-human · unremarkable · no-reference), separation and separationParaphrased (the share of machine text this statistic separated in its worst domain, before and after a paraphrase attack), reference (human and machine interquartile ranges, counts and histograms), and hits and spans for signals that match phrases. There is deliberately no aggregate score.

POST/detectauth optional · 200,000 characters

text (required).

Response. available and reason; abstained (true below minWords, 250); wordCount; probability (machine-likeness, calibrated); threshold (the probability at which it flags, set for a target false-positive rate); band (low · moderate · high · very_high); verdict (consistent_with_machine_generation or consistent_with_human_writing, never “written by AI”); measuredFpr (the false-positive rate measured on held-out human text for this model version); modelVersion; caveat. Paraphrased or “humanised” text drifts toward the machine ranges; on such text the number describes the tool, not the writer.

POST/reportauth optional

text (required), password. Returns { report, signature }. report holds version, issuedAt, detector (name, endpoint, frame, ecc), subject (sha256, bytes, characters of the submitted text; the text itself is not included), result (marked, authenticated, conflict, candidateCount, technique, watermark, techniquesDetected) and caveats. signature holds alg (Ed25519), publicKey (base64 SPKI), keySource (configured or ephemeral-process-key) and value (base64). The signature covers the canonical JSON of report; the specification gives the canonical form and a verifier. This endpoint attests to the TMK2 frame only; use /provenance for the cross-scheme view.

GET/report/key

Returns alg, publicKey, keySource and reportVersion.

5. Accounts, keys, history, usage

Accounts exist so that an integration can have history and counters. They are free, and nothing below is needed to watermark or inspect text.

GET/status

Returns configured (accounts available) and message.

POST/auth/register

email, password (8 characters or more), displayName. Returns 201 with user (id, email), session (access_token, refresh_token, expires_in, token_type) or null when email confirmation is pending, and message.

POST/auth/login

email, password. Returns access_token, refresh_token, expires_in, token_type and user. Use the access token as the bearer for session endpoints.

POST/auth/logoutsession

Returns { ok: true }. The client discards its token; expiry does the rest.

GET/auth/mesession or API key

Returns id, email, kind (jwt or api_key) and profile (id, display_name, created_at) or null.

POST/keyssession

name (required). Returns 201 with id, name, key_prefix, created_at, key (the full tm_… key, shown once) and warning. Only a SHA-256 of the key is stored.

GET/keyssession

Returns keys: id, name, key_prefix, created_at, last_used_at, revoked_at.

DELETE/keys/:idsession

Revokes the key. Returns key with revoked_at set; 404 if it is not yours or already revoked.

GET/watermarks · /watermarks/:idsession or API key

Returns watermarks (newest first, at most 100) or one watermark: id, technique, authenticated, cover_preview, payload_preview, frame_bytes, channels, created_at. Previews only; the full cover and payload are never stored.

GET/statssession or API key

Returns configured, windowDays (30), since, totals (calls, errors, successRate, avgMs), byEndpoint (calls, errors, avgMs per path, with ids collapsed to :id), byDay, and instance – counters for the one server process that answered, which reset on every restart and are never folded into the totals.

6. Changes

  • 3 September 2026. CORS opened with preflight support. Anonymous rate limiting (429 + Retry-After); credentialed requests exempt. Text fields capped at 500,000 characters (413). No response shape changed.
  • August 2026. /provenance gained the vendor tier and the StegCloak, SNOW, unidentified-carrier and caller-key rows, additively. /evidence and /detect added. /stats added.

AI Text Mark is a free tool provided by ROGA AI LIMITED. These terms set out what you can expect from it, what it does not promise, and the rules for using it.

Effective 11 August 2026

1. Who provides this service

AI Text Mark ("the Service") is operated by ROGA AI LIMITED, a company registered in Gibraltar under number 125994, with its registered office at Unit G02, Eurocity, Europort Avenue, Gibraltar, GX11 1AA ("we", "us", "our"). Contact: legal@rogaai.com.

2. Accepting these terms

By using the Service, including the web interface and the API, you agree to these terms. If you do not agree, do not use it. If you use the Service for an organisation, you confirm you are authorised to accept these terms on its behalf. You must be at least 18, or the age of majority where you live.

3. The Service is free

There is no charge, no subscription and no paid tier. Because nothing is paid for, nothing is owed in return: we do not commit to availability, uptime, response times, support, or continued existence of any feature. We may change, limit, suspend or discontinue the Service, in whole or in part, at any time and without notice.

We may apply rate limits or other restrictions to keep the Service usable for everyone. Sponsorship shown in the interface does not entitle a sponsor to influence how the Service works or what these terms say.

4. What the Service actually does, and what it does not

This section matters more than the rest, so it is written plainly rather than in legal shorthand.

The Service embeds a machine-readable mark into text you supply, and reads such marks back out. It is a post-hoc, format-based watermark: the mark lives in the character encoding, not in the words.

Do not rely on the Service for any of the following
  • Robustness. Marks are removable. Retyping, OCR, screenshots, printing and rescanning, paraphrasing, and ordinary text normalisation destroy them. So does a single well-chosen regular expression. We publish these failure modes rather than hide them, and you accept them as a known property of the Service.
  • Proof of authorship or origin. A mark that carries no verified authenticator can be forged by anyone who understands the format, and can be lifted out of one document and placed into another. Only an authenticated result carries any weight, and even then it evidences that a holder of the password produced that payload, not that any particular person wrote the surrounding text.
  • Detection of other systems’ watermarks. The Service cannot tell you which AI model produced a piece of text. Where it reports that a scheme is not readable here, that is a statement about our capability, not evidence that the text is unmarked.
  • Absence of a mark. A negative result never proves text is unmarked or human-written.
  • Regulatory compliance. Nothing in the Service constitutes legal advice or a compliance determination. Using it does not make you or your organisation compliant with the EU AI Act or any other law. Obligations such as Article 50(2) of Regulation (EU) 2024/1689 fall on the provider of the generative system, and cannot be discharged by using a third-party tool. Signed detection reports record what our detector observed in the bytes submitted at the time stated, and nothing more.

Do not use the Service to make accusations about individuals. Treating any output as evidence of academic dishonesty, fabrication or misconduct is outside what the Service supports, and you accept sole responsibility for doing so.

5. Your content

You keep all rights in the text you submit. You grant us only the limited permission needed to process it and return a result.

We do not store the text you submit. Watermark history, where you are signed in, records short previews and metadata only, never the full cover text or payload. Usage counters record the endpoint, the day, call counts and timing, and nothing about the content of requests. You are responsible for having the right to submit whatever you submit, and for not submitting personal data, confidential information or regulated data that you are not permitted to process this way.

6. Accounts and API keys

Accounts and API keys are optional; the core tools work anonymously. If you create one, keep your credentials secure and do not share them. An API key is shown once and is your responsibility from that moment. You are accountable for activity under your account or key. Tell us promptly at legal@rogaai.com if you believe either has been compromised.

7. Acceptable use

Do not use the Service to:

  • break the law, or infringe anyone’s rights;
  • mark text in order to impersonate a person or organisation, or to make a false claim of origin;
  • remove or alter a watermark in content you have no right to modify, or to defeat another party’s provenance or licensing controls;
  • embed malicious payloads, or use invisible characters to smuggle instructions into another system, including prompt injection against AI systems;
  • submit material you have no right to process, or personal data in breach of applicable data protection law;
  • attack, overload, probe or circumvent the Service, its rate limits or its authentication;
  • present output as a compliance certification, an official approval, or forensic evidence.

We may suspend or remove access at any time, without notice, where we reasonably believe these rules have been broken.

8. Intellectual property

The Service, its interface, documentation and branding remain ours or our licensors’. These terms grant you no rights in them beyond using the Service as intended. The frame format is published so that others can read marks independently; publishing a format is not a transfer of any trade mark or other right.

9. No warranty

The Service is provided “as is” and “as available”, without warranty of any kind, whether express, implied or statutory, to the fullest extent the law allows. We do not warrant that it will be uninterrupted, secure, error-free, or fit for any particular purpose, nor that any mark will be embedded, survive, or be recovered in any given case.

10. Limitation of liability

To the fullest extent permitted by law, we are not liable for any indirect, incidental, special, consequential or punitive damages, nor for loss of profits, revenue, data, goodwill or business, arising from or connected to your use of the Service.

Because the Service is provided free of charge, our total aggregate liability to you for all claims is limited to one hundred pounds sterling (£100).

Nothing in these terms limits liability that cannot lawfully be limited, including liability for death or personal injury caused by negligence, or for fraud or fraudulent misrepresentation. If you are a consumer, these terms do not affect your statutory rights.

11. Indemnity

You agree to indemnify and hold us harmless against claims, losses and reasonable costs arising from your use of the Service, your content, or your breach of these terms – in particular any claim arising from how you characterised or relied upon the Service’s output.

12. Suspension and termination

You may stop using the Service at any time, and may delete your account if you have one. We may suspend or terminate access at any time, with or without cause or notice. On termination, sections that by their nature should survive – including sections 4, 9, 10 and 11 – continue to apply.

13. Changes to these terms

We may update these terms. The effective date at the top will change, and continuing to use the Service after that means you accept the revised version. If a change is material, we will make reasonable efforts to signal it in the interface.

14. Governing law

These terms are governed by the laws of Gibraltar, and the courts of Gibraltar have exclusive jurisdiction, except where mandatory law in your country of residence gives you the right to bring proceedings elsewhere. If any provision is held unenforceable, the rest continues in force.

15. Contact

ROGA AI LIMITED, Unit G02, Eurocity, Europort Avenue, Gibraltar, GX11 1AA. Company number 125994. Questions about these terms: legal@rogaai.com.

What this service stores, what it does not, and who else sees anything. The short version: the text you submit is never stored, and there are no cookies.

Effective 11 August 2026

1. Who is responsible

ROGA AI LIMITED, registered in Gibraltar under number 125994, registered office Unit G02, Eurocity, Europort Avenue, Gibraltar, GX11 1AA, is the controller for personal data described here. Contact: legal@rogaai.com.

2. The text you submit is not stored

Cover text, watermark payloads and anything you paste for extraction or analysis are processed in memory to produce a result and then discarded. They are not written to a database, not logged, and not used to train anything.

If you are signed in, watermark history keeps a short preview of the cover and payload – a truncated snippet for your own reference – plus the technique, channel list and frame size. That is the only trace of a marked document we retain, and you can delete it.

Detection reports are generated and signed on request and returned to you. We keep a SHA-256 of the submitted text inside the report you download; we do not retain a copy server-side.

3. What we do store

DataWhyKept
Email address, and display name if givenTo create and authenticate your accountUntil you delete the account
API key name, prefix, and a SHA-256 hash of the keyTo authenticate API calls. The key itself is shown once and never storedUntil you revoke it or delete the account
Watermark history previews and metadataSo you can see what you markedUntil you delete it or the account
Usage counts per endpoint per day, with timingTo show you your own usage and keep the service runningRolled up daily; deleted with the account

Usage counters record the endpoint, the date, a call count, an error count and total latency. They record nothing about the content of requests, and no IP addresses.

Accounts are optional. The embed, extract, analyze, compare and provenance tools all work without one, and anonymous use is counted only in aggregate against a shared placeholder identifier.

4. Cookies and browser storage

This site sets no cookies. Two things are kept in your own browser and never sent anywhere except as described:

  • textmark_access_token in session storage, when you sign in. It is your session and disappears when you close the tab.
  • textmark.mode and textmark.consent in local storage – your Simple/Expert preference and your analytics choice.

Both are strictly necessary for the tool to function or to honour a choice you made, so neither is subject to consent. You can clear them at any time through your browser.

5. Analytics, and your choice

We use Vercel Web Analytics to count page views. It is cookieless and does not build a profile of you or follow you across sites. It is still optional, so it does not load at all unless you allow it – the notice you saw on arrival gates the script rather than merely announcing it.

6. Who else processes data

  • Supabase – authentication and database hosting, for accounts, API keys, history and usage counters.
  • Vercel – hosting and delivery. Vercel processes request metadata, including IP addresses, as part of serving and protecting the site, and provides the optional analytics above.

Typefaces are served from this site rather than from a font host, so simply loading the interface sends your IP address to no one but our hosting provider.

We do not sell personal data, and we do not share it for advertising.

7. Legal basis and international transfers

Where the UK GDPR or EU GDPR applies, we rely on: contract, for running an account you asked for; legitimate interests, for keeping the service secure and for aggregate usage counts; and consent, for optional analytics, which you may withdraw at any time.

Our processors may store or process data outside Gibraltar, including in the EU, the UK and the United States, under the transfer safeguards those providers offer.

8. Your rights

Subject to applicable law, you may request access to your personal data, correction, deletion, a portable copy, or restriction of processing, and you may object to processing based on legitimate interests. Deleting your account removes your profile, API keys, watermark history and usage counters.

Write to legal@rogaai.com. You also have the right to complain to your local data protection authority.

9. Security, children, and changes

API keys are stored only as SHA-256 hashes, passwords are handled by our authentication provider and never seen by us, and database access is restricted per user. No system is perfectly secure, and we do not claim otherwise.

The service is not intended for children under 16, and we do not knowingly collect their data.

We may update this notice; the effective date above will change, and material changes will be signalled in the interface.

Nothing to hand? Try a real one:

0 words

Paste text on the left, or pick a sample, then press Analyze text.

You get two separate answers that are never merged into one score. The classifier is a statistical model trained on human and machine text: it abstains below 250 words and reports a calibrated machine-likeness with our measured false-positive rate. The evidence rows are twelve hand-checkable writing statistics, each placed against real human writing, with what it catches and where it misleads.

Neither is proof of authorship. No text tool can offer that.