AI data leak detection with honeytokens

When spam or a phishing attempt arrives, most companies can only guess where their data escaped. A canary agent removes the guessing: it gives every vendor its own bait address, then watches those addresses for as long as it takes.

Can AI detect which vendor leaked my data?

Yes, if you plant the evidence in advance. An AI agent can mint a unique email alias on your own domain for each vendor you share data with, seed a matching fake contact record into that one vendor's system, and watch those addresses continuously. Mail arriving at a single-vendor alias points at that vendor. Qoren runs the watching always-on, and the escalation email is drafted for a person to approve.

The problem

Companies hand contact data to dozens of tools and partners: a CRM, an email platform, a form builder, a lead list vendor, an agency. When that data turns up somewhere it should not be, there is no way to tell which of the dozens was responsible, so nobody is ever held to the contract they signed. The trick that solves it is old and simple, unique bait per vendor, but it only works if someone keeps the register current and keeps watching the bait for years. That is exactly the kind of patient, boring vigilance a person will not sustain and an always-on agent will.

How an agent handles it

  • Mint a unique email alias on your own domain each time you sign up for a tool, upload a list, or share data with a partner.
  • Seed a plausible but fake contact record, matched to that alias, into that one vendor's system and nowhere else.
  • Keep a register mapping every alias and bait record to the vendor, the date, and what was shared.
  • Watch the bait inboxes continuously, for as long as the relationship and the risk last.
  • Capture the evidence on a hit: sender, headers, timestamps, and the alias that was used.
  • Draft the notification or data processing agreement query for a person to review and send.

Why it sells

  • Attribution instead of suspicion the first time bait mail arrives.
  • A vendor trust ledger that gets more valuable the longer it runs.
  • Dated evidence ready to bring into a renewal or renegotiation conversation.

How it runs, step by step

Data leak canary6
  1. Mint the bait

    Every time the company onboards a vendor, uploads a contact list, or shares data with a partner, the agent mints a unique alias on your own domain, something like acme-crm-7f2@yourdomain, and writes a plausible but entirely fictional contact row to go with it: a name, a title, a phone number that belongs to nobody. That bait goes into that one vendor's system and nowhere else, which is what makes any future hit attributable.

  2. Keep the register

    The mapping is the asset, not the alias. The agent maintains a register of every alias and bait record: which vendor received it, on what date, through which upload or signup, what fields it contained, and which contract or data processing agreement governs it. Aliases minted years apart stay legible because the register is written at the moment of seeding, when the context is still known.

  3. Watch, indefinitely

    The agent runs always-on and keeps every bait address under watch on a schedule. This is the part that fails when a human owns it, because the payoff may be two years out and the daily result is nothing. For an agent, a watch that finds nothing thousands of times in a row costs almost nothing to keep running, and the value only accrues with time.

  4. A hit arrives

    The day mail lands at a single-vendor alias, the agent captures the evidence rather than just alerting: the full headers, sending domain and IP, the timestamp, the content, and the register entry that says which vendor holds that alias and since when. It classifies the hit as marketing mail, list resale, or phishing, and notes whether other aliases from the same vendor were touched.

  5. The drafted escalation

    The agent then drafts the outbound message, addressed to the vendor's privacy or security contact, laying out the alias, the seeding date, the receipt date, and the attached evidence. It arrives in draft for a person to read, edit, and send. Attribution is evidence, not a legal finding, so the decision about what to assert and what to demand stays with a human, and with counsel where it belongs.

  6. The vendor trust ledger

    On a recurring schedule the agent reports the whole picture: which vendors hold bait, how long each has been clean, which produced hits and of what kind. That ledger is a real input to vendor management, so a leak discovered in March is still sitting in the renewal brief in November instead of being forgotten.

When this is not the fit

This is a canary, not a data loss prevention suite or a forensics practice. It cannot see inside a vendor's systems, will not catch exfiltration through channels you did not bait, and says nothing about data you shared before the seeding started. A hit tells you an address only that vendor held ended up in someone else's hands, which is strong evidence and still not proof of fault: the vendor may have been breached rather than careless, and the record may be ambiguous. Regulated incident response, breach notification duties, and anything heading toward a dispute belong with your security and legal professionals, with the agent supplying the timeline rather than the judgment.

Frequently asked questions

How does the agent know which vendor leaked the data?

By elimination, planned in advance. Each alias and bait record is given to exactly one vendor and recorded in the register, so mail arriving at that alias could only have come from data that vendor held. That is attribution, and it is evidence worth acting on, but it is not a legal finding on its own: the vendor may have sold the list, been breached, or passed it to a subprocessor. Deciding what it means stays human.

What about false positives?

They happen, and the agent is built to sort them. Legitimate mail from the vendor itself is expected and gets logged rather than escalated. Guessed addresses are unlikely because aliases are random rather than predictable, and the agent notes corroborating signals such as several aliases from one vendor being hit in the same window. Anything unclear goes into the report instead of triggering an accusation, and nothing goes out without a person approving it.

Is seeding honeytokens legal and ethical?

The setup deliberately keeps everything inside your own perimeter: aliases on your own domain, in accounts your company owns, holding contact records that describe nobody real. No third party is deceived about anything material, no real person's data is used as bait, and no system is accessed that is not yours. It is closer to marking your own property than to surveillance. Even so, check the arrangement against your own vendor agreements and your counsel's advice before you roll it out.

What is the vendor trust ledger?

It is the report that accumulates from all this patience: every vendor holding bait, how long each has stayed clean, and every hit with its date and classification. It turns a vague sense that a vendor is sloppy into a dated record, which is a much better thing to have on the table at renewal time or when the security review comes around.

Deploy this for your clients.

Pick a template, configure it per client, and go live in minutes. Plans from $39/mo.