How EnfoldAI Spots a Scam Message: Every SMS Is Attacker-Written
A single e-challan scam, followed from the system broadcast that delivers it to the moment someone decides whether to tap the link in it — and the seven invariants an internal audit taught us to write down.
Mayank TalwarCo-founder and CTO
Krishna AgrawalFounding Engineer
Gaurav AgarwalFounding Engineer
6 October 202613 min read
+91 98XXX XXX21 · Today, 9:41 pm
Your vehicle has a pending e-challan of Rs.1,000. Pay before 11:59 pm today to avoid court action: echallan-parivahan.example/pay
Links offPending — not checked yet
Scam messages like this one are a routine part of India's SMS traffic, and they work. In 2025, Indians reported 28.15 lakh cyber-fraud cases and lost ₹22,495 crore, according to Ministry of Home Affairs figures reported by ThePrint.
EnfoldAI can run as an Android phone's default SMS app, so every incoming message passes through our code before its recipient sees it. That's a privileged position, and it means we have to deliver every message, including the OTP someone is waiting for at a checkout, even when our backend is struggling.
This post follows a single scam message from arrival to the moment someone decides whether to tap a link in it. Three ideas shaped most of our decisions. Every byte of an SMS is written by whoever sent it, including the parts that look like metadata. Checking who sent a message is cheaper and more reliable than judging what it says. And an unknown verdict is not a safe verdict, so the code that handles a missing answer matters as much as the model.
Background and threat model
An Indian inbox gets application-to-person (A2P) messages, like bank alerts and OTPs, from businesses, and person-to-person messages from ordinary 10-digit numbers. A2P senders have to register their sender IDs (called DLT headers) and message templates on the operators' Distributed Ledger Technology (DLT) platforms under TRAI's rules. Since a February 2025 amendment, every header also ends in a suffix that declares its type, such as -T for transactional or -G for government. That registry is the closest thing we have to ground truth about who sent a message, and scammers can't write to it.
We treat every input on the path as untrusted until we have checked it, and the table below shows how far we trust each one.
| Input | Who controls it | How we treat it |
|---|---|---|
| Message body: text, links, Unicode | The sender, completely | Hostile, and fenced off whenever a model reads it |
| Sender address or header | The sender, within network limits | Checked against the registry; an alphanumeric sender never matches a contact |
| Our backend's verdict | Us, but it can be late or missing | Semi-trusted, and parsed against an allowlist on the phone |
| Registry data | The operators | Trusted, but changes are capped and sensitive names are reviewed by a person |
Invariants
Before we picked a model, we wrote down invariants that had to hold whichever model, rule or server was misbehaving. The rest of the post refers back to them.
| ID | Invariant | Enforced by |
|---|---|---|
| I1 | A message reaches the inbox without waiting for a verdict | Deliver-first ingest on the phone |
| I2 | Unknown never renders as safe | One allowlist parser, and links stay off until cleared |
| I3 | Every verdict comes from exactly one rule, and a high-risk verdict says which | An ordered rule cascade |
| I4 | No model decides alone, and a failed second-opinion call never weakens a verdict | Models label, rules decide; only an explicit "yes" counts |
| I5 | Nothing is uploaded for scoring without consent and sign-in | A consent gate at every upload point |
| I6 | Anonymous requests never reach the second-opinion or OTP-extraction calls | A frozen per-request object, checked by a test |
| I7 | Policy changes ship without an app release | Rules, lists, registry and labels live on the server |
Architecture overview

Here's what happens to one message as it moves through Figure 1.
- Ingest. Android delivers the message through
SMS_DELIVER, a broadcast only the system can send. We write it to the inbox straight away, record it as pending in a local Room database with its links disabled, and hand off to a foreground service (or an expedited WorkManager job if the service can't start). - Trust boundary 1. If the sender isn't a strictly matched or blocked contact, and the user has consented and signed in, a pinned, authenticated client sends the text, the sender, a contact flag and a device identifier to our API.
- Scoring. An input guard authenticates the caller, caps input sizes and enforces usage budgets. Sender identity, link analysis and a classifier produce evidence, an ordered set of rules turns it into a verdict, and high-risk results are recorded for human review.
- Trust boundary 2. The verdict returns to the phone as semi-trusted input, where a strict parser maps it to a state and presentation rules decide what the user can see and tap.
None of this was designed in one go. The first version, in 2024, was a set of hand-written rules, and over time we added curated sender lists, tripwire rules, first-run inbox scoring, LLM labelling and then our own classifier. In September 2026 an internal security audit found places where the system failed open, quietly treating an answer it did not recognise as good news. Most of the invariants above came out of that audit.
Sender identity: a registry mirror guarded like code

We resolve identity first because it costs a single lookup and the sender cannot forge it. The parser treats the operator prefix and type suffix as optional and looks up the bare header by exact match. Every sender also gets a category, such as a registered header or a 10-digit number, which the classifier later sees.
There's no single API for the registry. Each operator publishes its own list, and six of the seven we read are PDFs (the other is a CSV). A scheduled job downloads them with file-signature, size and same-host redirect checks, parses the PDF rows by serial number and diffs the result against the last snapshot, publishing only after a clean, complete run.

The guards are there because the mirror changes outcomes as well as answering lookups. A place in the registry earns a sender extra trust, so one poisoned row could make phishing from that header look trusted for every user. Closing that risk was the point of the pull request that hardened this job. Large batches of deletions, additions or renames now wait for someone to release them, and a rename that touches a sensitive kind of sender is quarantined until a person approves that exact header and name. Headers that look registered but are missing from the mirror go to a review queue.
This has held up well. The gaps are freshness, since the mirror is only as current as its last run, and person-to-person numbers, which have no registry at all, so identity can raise our suspicion of them but never lower it.
Content understanding: telling a scam message from an OTP
To work out what a message is trying to do, we use an in-house, base-size multilingual transformer encoder that we fine-tuned to give each SMS one of fifteen labels. Five are legitimate categories, such as OTP and transaction, and ten are scam tactics, such as impersonation and manufactured urgency. It never sees the text on its own.

Ahead of the message we add fields we compute ourselves, namely the sender's category and header, the registered business name, a profile of any link it carries and a flag for look-alike or invisible characters. The only part a scammer writes is the message. We also strip the "suspected spam" prefixes some carriers add, which would otherwise teach the model to read the carrier's verdict.
We tune the model toward precision. Verified senders get more benefit of the doubt than unknown ones, and sender and link evidence settle the cases the text alone can't. The model runs on CPUs, and answers are cached by message and sender, so a campaign blast is classified once.
Links get read twice. A strict extractor picks out well-formed URLs with bounded-repetition patterns, and a looser detector flags anything that merely resembles a link, including obfuscated and encoded addresses. The two are compared host by host, and if the loose pass finds a host the strict one missed, the message loses the benefit of the doubt.
The scoring path never fetches the page behind a link, though, and language coverage is only as good as the training data.
The LLM as a fenced specialist
LLMs read messy, multilingual text well, but they also follow instructions hidden in text, which is how prompt injection works, and a scam message is written to make someone act. We use them for three narrow jobs only: standing in when the classifier is down (with the same fifteen labels), extracting an OTP our own extractor missed, and giving a yes-or-no second opinion, which is the only one that can soften a verdict.
That second opinion is reserved for a narrow set of borderline cases, where the rules would flag what may well be an ordinary business message, and only where a wrong "yes" would cost the user little. In the pull request, the engineer who drew those limits put the principle as "a signal the attacker can type is not evidence".

The rubric and a registry fact we compute go in as system messages the sender can't write into, and the SMS comes last, delimited, length-capped and labelled as data rather than instructions.
# Simplified for this post. Real prompts, schemas, names and limits differ.
def second_opinion(sender: str, body: str, registry_fact: str) -> bool:
messages = [
{"role": "system", "content": RUBRIC}, # what "legitimate" means
{"role": "system", "content": registry_fact}, # computed by us; the sender cannot write here
{"role": "user", "content": fence(sender, body[:MAX_CHARS])}, # data to read, not obey
]
try:
args = ask_with_retries(messages, tool=VERDICT_TOOL) # forced function call
return args.get("legitimate") is True # missing, null or the string "true" all mean no
except Exception: # timeout, no tool call, malformed output
return False # failure changes nothing, the finding stands
Only a literal boolean true counts, so a missing field, a string, an error or a timeout leaves the high-risk finding in place (I4). The same PR described the design as "the model confirms, it does not decide."
OTP extraction deserves a plain answer, because a one-time code is the most sensitive thing in an inbox. Our own extractor always runs first. Only when it finds no code in a message already labelled as an OTP do we send that message's text to a language model, with no sender, account or device details attached, and only for signed-in users who have agreed to scanning.
Cost is handled structurally too. Each request gets a frozen object holding the two paid calls, and anonymous requests get no-spend stand-ins. As the PR explained, "a flag has to be remembered; a bound object does not." A test checks this (I6) by running anonymous scoring with the real calls patched to raise.
The second opinion can still be wrong where it's allowed to soften a verdict, which we accept only because we limit it to cases where a wrong answer is low-cost.
From signals to a verdict: the ordered cascade

The verdict comes from an ordered list of rules, where the first match decides and a high-risk rule always names itself (I3). Model labels and the second opinion are evidence that individual rules read, so no model decides a verdict on its own (I4).
We did consider a weighted score. A score tells you how bad something looks, but the cascade also tells you why, which is what users and analysts need, and policy changes become a matter of ordering reviewed, tested rules rather than tuning weights. A refusal never masquerades as a verdict either, so a caller over budget gets an HTTP 429 instead (I2).
The trade-off is that rule order is policy, and a new rule can quietly make a later one unreachable, so reordering gets the same scrutiny as code.
The last mile: verdicts on the phone
Plenty of writing about scam detection stops when the model answers. For us that's about halfway, because the phone still decides what the answer, or the lack of one, means for what a person can tap.

Every message is in exactly one state. Unknown is reserved for the server to say it cannot classify something, and the phone treats it as final, because retrying would cause what the PR author called "a retry storm aimed at a service that is already down". Answers from the live call, the retry sweep and the first-run backfill all go through the same parser.
// Simplified for this post.
enum class Verdict { HIGH, LOW, UNKNOWN }
/** A verdict we recognise, or null for "no answer". Never a guess. */
fun parseVerdict(raw: String?): Verdict? =
Verdict.entries.firstOrNull { it.name == raw?.trim()?.uppercase() }
/** Links are earned. Only a saved contact or a low-risk verdict turns them on. */
fun linksEnabled(verdict: Verdict?, savedContact: Boolean, hiddenChars: Boolean): Boolean =
!hiddenChars && (savedContact || verdict == Verdict.LOW)
Both exist because of a whole class of bug. Earlier versions checked for the one bad answer and treated everything else as clean, so an answer the phone didn't recognise read as safe. The September audit flagged it, and since then a message has had to earn its links.
| State | In the conversation | Links | Notification |
|---|---|---|---|
| Pending | Shown | Off | Once the scoring call returns or fails |
| Saved contact | Shown | On | Yes |
| Low risk | Shown, with its category label | On | Yes |
| Unknown | Shown | Off | Yes |
| High risk | Hidden: "Message hidden for your safety" | Off | None, and a late verdict replaces earlier ones with a warning that leaves out the text |
| Failed check | Shown as "Not checked" | Off | Only the one from arrival |
The rest of the phone-side design follows the same idea.
- Hidden characters. A display sanitizer removes zero-width, direction-override and similar characters. Zero-width joiners stay because Devanagari and emoji need them, but their presence turns links off, and scoring and reporting always use the original text.
- Notifications. Release 1.1.51 stopped holding back notifications for unscored messages, because OTPs never arrived when the network was down. A notification now waits only for the scoring call itself. A high-risk verdict means no notification, and anything else, including a failed call, notifies with cleaned text and a Copy OTP button. If a retry later finds a message high risk, the app clears that sender's notifications and posts a warning without the message text.
- Retries. We never retry a POST in the network layer, since a timeout doesn't prove the request failed. A WorkManager sweep retries while the phone is online instead, with conditional writes that only touch rows still pending.
- Contacts. Only a sender that could really be a phone number can match a saved contact, so a business-style sender ID can never borrow a friend's trust.
- Consent. Every call carrying message data checks for consent and sign-in (I5), and without them messages are still delivered but stay pending.
Each rule lives in a small, pure policy object tested in CI. The verdict parser has 6 tests, the live-scoring callback 9, the link policy 8, the sanitizer 13 and the retry sweep 22, and that coverage is why we can say the app fails closed.
There are still gaps. An offline phone leaves a message pending, with its links off and no verdict. Messages from saved contacts are also not scored by default, so a scam sent from a friend's compromised number is trusted unless the user turns on scanning for personal chats.
Humans in the loop
Our analysts work through three queues in the admin portal, covering high-risk messages to verify, unknown headers to promote and registry renames to approve. People can also forward suspicious messages, screenshots and PDFs to our WhatsApp line, where an AI drafts a reply and an analyst edits and sends it, so no AI-written verdict reaches anyone there unchecked.
Design considerations
| Pillar | How it shows up |
|---|---|
| Security | An explicit threat model, certificate pinning with a per-hop host guard, fenced LLM channels and bounded regular expressions |
| Privacy | Consent gates (I5), the narrow OTP path described above, and a CI check that fails the build when a log call is handed message text, an OTP or a URL |
| Reliability | Delivery first (I1), idempotent retries, a timeout on every upstream call, and refusals never mistaken for verdicts (I2) |
| Explainability | One rule per verdict (I3), and an app that shows "Not checked" instead of staying silent |
| Cost | An in-house model on CPUs doing most of the work, cached answers, usage budgets and no second-opinion or OTP calls for anonymous requests (I6) |
Lessons learned
- The classifier was the easy part. What the phone does with an answer, or without one, decides whether anyone is protected.
- Parse values that cross a trust boundary against an allowlist, including your own server's answers, so anything you don't recognise means "no answer". Both fail-open bugs here checked for the one bad value and treated everything else as safe.
- Keep anything an attacker can type away from the decision, and keep facts in channels they cannot write to.
- Make promises structural, because a test can check an object's structure but nobody reliably remembers a flag.
- Protect the data you trust most, since a sender registry is exactly what an attacker would want to poison.
Conclusion
Detecting a scam message is a classification problem, but protecting the person who receives it is a systems problem that depends on who sent the message, what a model may decide and what happens when an answer is late, wrong or missing. Most of it is unglamorous engineering, and little of it is specific to SMS. In upcoming posts we'll cover reading forwarded screenshots and PDFs without letting them instruct a model, and the registry pipeline.
If you're holding the phone
To spot a fake SMS, look at what it asks you to do. If a message pushes you to act within minutes, to pay, to call a number it gives you or to install something, treat it as a scam message until you've checked it some other way. Never share an OTP with anyone who asks for it, whoever they claim to be. If a link in EnfoldAI won't open, that's on purpose, because we haven't cleared it yet. If you've already lost money, call 1930, India's cyber-fraud helpline, or report it at cybercrime.gov.in as soon as you can.
What we leave out
We're happy to explain how our defences work, but not how to get around them. That's why this post leaves out thresholds, the rules themselves and their order, the exact conditions that trigger each check, model and vendor names, and hosting details. None of our protection depends on keeping the design secret, but there's no reason to hand out a tuning manual either.
Sources
- Ministry of Home Affairs cyber-fraud figures for 2025, reported by ThePrint, 21 February 2026.
- Telecom Regulatory Authority of India, Telecom Commercial Communications Customer Preference (Second Amendment) Regulations, 2025, 12 February 2025.
- National Cyber Crime Reporting Portal, cybercrime.gov.in.
Acknowledgments
This system is the work of EnfoldAI's Android, backend and data teams, and of the analysts who review what it flags. We especially thank our founding engineers, Krishna Agrawal and Gaurav Agarwal.