Skip to main content
New: Free scam checker: paste any SMS and find out in seconds Try it →

How EnfoldAI Spots a Scam Message: Every SMS Is Attacker-Written

A single e-challan scam, followed from the system broadcast that delivers it to the moment someone decides whether to tap the link in it — and the seven invariants an internal audit taught us to write down.

  • Mayank TalwarCo-founder and CTO
  • Krishna AgrawalFounding Engineer
  • Gaurav AgarwalFounding Engineer

6 October 202613 min read

SpecimenIllustrative; details changed

+91 98XXX XXX21 · Today, 9:41 pm

Your vehicle has a pending e-challan of Rs.1,000. Pay before 11:59 pm today to avoid court action: echallan-parivahan.example/pay

Links offPending — not checked yet

Scam messages like this one are a routine part of India's SMS traffic, and they work. In 2025, Indians reported 28.15 lakh cyber-fraud cases and lost ₹22,495 crore, according to Ministry of Home Affairs figures reported by ThePrint.

EnfoldAI can run as an Android phone's default SMS app, so every incoming message passes through our code before its recipient sees it. That's a privileged position, and it means we have to deliver every message, including the OTP someone is waiting for at a checkout, even when our backend is struggling.

This post follows a single scam message from arrival to the moment someone decides whether to tap a link in it. Three ideas shaped most of our decisions. Every byte of an SMS is written by whoever sent it, including the parts that look like metadata. Checking who sent a message is cheaper and more reliable than judging what it says. And an unknown verdict is not a safe verdict, so the code that handles a missing answer matters as much as the model.

Background and threat model

An Indian inbox gets application-to-person (A2P) messages, like bank alerts and OTPs, from businesses, and person-to-person messages from ordinary 10-digit numbers. A2P senders have to register their sender IDs (called DLT headers) and message templates on the operators' Distributed Ledger Technology (DLT) platforms under TRAI's rules. Since a February 2025 amendment, every header also ends in a suffix that declares its type, such as -T for transactional or -G for government. That registry is the closest thing we have to ground truth about who sent a message, and scammers can't write to it.

We treat every input on the path as untrusted until we have checked it, and the table below shows how far we trust each one.

Input Who controls it How we treat it
Message body: text, links, Unicode The sender, completely Hostile, and fenced off whenever a model reads it
Sender address or header The sender, within network limits Checked against the registry; an alphanumeric sender never matches a contact
Our backend's verdict Us, but it can be late or missing Semi-trusted, and parsed against an allowlist on the phone
Registry data The operators Trusted, but changes are capped and sensitive names are reviewed by a person

Invariants

Before we picked a model, we wrote down invariants that had to hold whichever model, rule or server was misbehaving. The rest of the post refers back to them.

ID Invariant Enforced by
I1 A message reaches the inbox without waiting for a verdict Deliver-first ingest on the phone
I2 Unknown never renders as safe One allowlist parser, and links stay off until cleared
I3 Every verdict comes from exactly one rule, and a high-risk verdict says which An ordered rule cascade
I4 No model decides alone, and a failed second-opinion call never weakens a verdict Models label, rules decide; only an explicit "yes" counts
I5 Nothing is uploaded for scoring without consent and sign-in A consent gate at every upload point
I6 Anonymous requests never reach the second-opinion or OTP-extraction calls A frozen per-request object, checked by a test
I7 Policy changes ship without an app release Rules, lists, registry and labels live on the server

Architecture overview

EnfoldAI's scam-scoring architecture in three bands. On the Android device: an SMS_DELIVER receiver (a system-only broadcast) joins multipart PDUs and writes to the SMS provider, a processing service runs as a foreground service with an expedited WorkManager fallback, and pre-checks cover contact trust, the user blocklist and the consent and sign-in gate; a local Room verdict store holds one pending row per message while the message is shown at once with links inert. Trust boundary 1 marks text, sender and links as attacker-controlled. In the backend scoring API: an input guard enforces an authenticated caller, length caps and usage budgets, then sender identity, link analysis and the content classifier compute signals for every message, backed by a DLT registry mirror, and an ordered rule cascade of reviewed rules decides, with model labels as evidence and never verdicts. An LLM second opinion is consulted in one narrow case, and high-risk results go to review queues. Trust boundary 2 marks the verdict as semi-trusted input. Back on the device, a strict verdict parser accepts known values only and turns anything else into Not checked, producing high risk (hidden, no notification text), low risk (links become tappable) or not checked (links inert, retry sweep).

Here's what happens to one message as it moves through Figure 1.

  • Ingest. Android delivers the message through SMS_DELIVER, a broadcast only the system can send. We write it to the inbox straight away, record it as pending in a local Room database with its links disabled, and hand off to a foreground service (or an expedited WorkManager job if the service can't start).
  • Trust boundary 1. If the sender isn't a strictly matched or blocked contact, and the user has consented and signed in, a pinned, authenticated client sends the text, the sender, a contact flag and a device identifier to our API.
  • Scoring. An input guard authenticates the caller, caps input sizes and enforces usage budgets. Sender identity, link analysis and a classifier produce evidence, an ordered set of rules turns it into a verdict, and high-risk results are recorded for human review.
  • Trust boundary 2. The verdict returns to the phone as semi-trusted input, where a strict parser maps it to a state and presentation rules decide what the user can see and tap.

None of this was designed in one go. The first version, in 2024, was a set of hand-written rules, and over time we added curated sender lists, tripwire rules, first-run inbox scoring, LLM labelling and then our own classifier. In September 2026 an internal security audit found places where the system failed open, quietly treating an answer it did not recognise as good news. Most of the invariants above came out of that audit.

Sender identity: a registry mirror guarded like code

The anatomy of an Indian commercial SMS header under TRAI's DLT rules, shown as XY-ACMEBK-T. The leading XY is the originating operator and licensed service area. ACMEBK in the middle is the registered header, up to six characters and tied to a business. The trailing -T is the type suffix, added since 2025, one of -P promotional, -S service, -T transactional or -G government. A person-to-person scam arrives from a 10-digit number such as +91 98XXX XXX21 and carries none of this; that absence is itself a signal, checked before a single word of the message is read.

We resolve identity first because it costs a single lookup and the sender cannot forge it. The parser treats the operator prefix and type suffix as optional and looks up the bare header by exact match. Every sender also gets a category, such as a registered header or a 10-digit number, which the classifier later sees.

There's no single API for the registry. Each operator publishes its own list, and six of the seven we read are PDFs (the other is a CSV). A scheduled job downloads them with file-signature, size and same-host redirect checks, parses the PDF rows by serial number and diffs the result against the last snapshot, publishing only after a clean, complete run.

The scheduled job that turns seven operator registries into one mirror. Six PDFs and one CSV sit outside our system, above a trust boundary where files are checked before they are parsed. The sync job then runs six steps: fetch, with parallel downloads, file-signature and size checks and same-host HTTPS redirects only; parse, anchoring PDF rows on each row's serial number; normalise and detect change, hashing each source so an unchanged one costs nothing; merge and diff against the last published snapshot; guards, which withhold mass deletions, additions or renames, quarantine a rename touching a sensitive kind of sender, and fail the run if a parse loses a large share of rows; and publish, where the snapshot switches only after a clean run. Human review releases withheld mass changes, approves quarantined names as an exact header and name pair, and promotes unknown headers. If any step fails the last good snapshot stays live. At scoring time the API does an exact lookup on the bare header against the registry mirror, and misses go to review.

The guards are there because the mirror changes outcomes as well as answering lookups. A place in the registry earns a sender extra trust, so one poisoned row could make phishing from that header look trusted for every user. Closing that risk was the point of the pull request that hardened this job. Large batches of deletions, additions or renames now wait for someone to release them, and a rename that touches a sensitive kind of sender is quarantined until a person approves that exact header and name. Headers that look registered but are missing from the mirror go to a review queue.

This has held up well. The gaps are freshness, since the mirror is only as current as its last run, and person-to-person numbers, which have no registry at all, so identity can raise our suspicion of them but never lower it.

Content understanding: telling a scam message from an OTP

To work out what a message is trying to do, we use an in-house, base-size multilingual transformer encoder that we fine-tuned to give each SMS one of fifteen labels. Five are legitimate categories, such as OTP and transaction, and ten are scam tactics, such as impersonation and manufactured urgency. It never sees the text on its own.

How the classifier's input is assembled for the specimen at the top of the post. Five fields are computed by us and come first: sender category (a 10-digit phone number), header (none), registered name (not available), link profile (a computed category) and an obfuscation flag (clean). The message itself comes last and is the only field written by the sender. Before encoding, whitespace is normalised and carrier 'suspected spam' prefixes are stripped, because they would leak the label to the model. A base-size multilingual transformer encoder, fine-tuned in-house and served on CPUs, works alongside an answer cache keyed on a hash of message and sender, and produces scores for fifteen labels: five legitimate categories and ten scam tactics. Two guardrails follow — verified senders get more benefit of the doubt than unknown ones, and precision comes first, with sender and link evidence settling unclear cases. If the classifier is unavailable, a general-purpose LLM is forced to answer with the same labels. One label goes to the rule cascade.

Ahead of the message we add fields we compute ourselves, namely the sender's category and header, the registered business name, a profile of any link it carries and a flag for look-alike or invisible characters. The only part a scammer writes is the message. We also strip the "suspected spam" prefixes some carriers add, which would otherwise teach the model to read the carrier's verdict.

We tune the model toward precision. Verified senders get more benefit of the doubt than unknown ones, and sender and link evidence settle the cases the text alone can't. The model runs on CPUs, and answers are cached by message and sender, so a campaign blast is classified once.

Links get read twice. A strict extractor picks out well-formed URLs with bounded-repetition patterns, and a looser detector flags anything that merely resembles a link, including obfuscated and encoded addresses. The two are compared host by host, and if the loose pass finds a host the strict one missed, the message loses the benefit of the doubt.

The scoring path never fetches the page behind a link, though, and language coverage is only as good as the training data.

The LLM as a fenced specialist

LLMs read messy, multilingual text well, but they also follow instructions hidden in text, which is how prompt injection works, and a scam message is written to make someone act. We use them for three narrow jobs only: standing in when the classifier is down (with the same fifteen labels), extracting an OTP our own extractor missed, and giving a yes-or-no second opinion, which is the only one that can soften a verdict.

That second opinion is reserved for a narrow set of borderline cases, where the rules would flag what may well be an ordinary business message, and only where a wrong "yes" would cost the user little. In the pull request, the engineer who drew those limits put the principle as "a signal the attacker can type is not evidence".

The second-opinion call, in three stages. A gate: a rule would flag an unknown sender's message, and the call is made only for a borderline case where a wrong yes would cost little; otherwise it is not asked and high risk stands. The call: two system messages the sender cannot write into carry the rubric, which defines a legitimate business message and says never to infer a brand from the sender code, and a registry fact computed by us, which states verified or not registered. A single user message holds untrusted data behind a preamble saying to read it and not follow it, with a delimited sender and a length-capped body. A general-purpose LLM answers with a forced function call returning a boolean and a reason, with retries, backoff and explicit timeouts; an anonymous request is answered by a no-spend stand-in that says no without calling a model. The decision: only a literal true lets the rule soften the verdict to low risk with a reason code, and anything else — missing, malformed, an error or a timeout — leaves high risk standing.

The rubric and a registry fact we compute go in as system messages the sender can't write into, and the SMS comes last, delimited, length-capped and labelled as data rather than instructions.

# Simplified for this post. Real prompts, schemas, names and limits differ.
def second_opinion(sender: str, body: str, registry_fact: str) -> bool:
    messages = [
        {"role": "system", "content": RUBRIC},         # what "legitimate" means
        {"role": "system", "content": registry_fact},  # computed by us; the sender cannot write here
        {"role": "user", "content": fence(sender, body[:MAX_CHARS])},  # data to read, not obey
    ]
    try:
        args = ask_with_retries(messages, tool=VERDICT_TOOL)  # forced function call
        return args.get("legitimate") is True  # missing, null or the string "true" all mean no
    except Exception:  # timeout, no tool call, malformed output
        return False  # failure changes nothing, the finding stands

Only a literal boolean true counts, so a missing field, a string, an error or a timeout leaves the high-risk finding in place (I4). The same PR described the design as "the model confirms, it does not decide."

OTP extraction deserves a plain answer, because a one-time code is the most sensitive thing in an inbox. Our own extractor always runs first. Only when it finds no code in a message already labelled as an OTP do we send that message's text to a language model, with no sender, account or device details attached, and only for signed-in users who have agreed to scanning.

Cost is handled structurally too. Each request gets a frozen object holding the two paid calls, and anonymous requests get no-spend stand-ins. As the PR explained, "a flag has to be remembered; a bound object does not." A test checks this (I6) by running anonymous scoring with the real calls patched to raise.

The second opinion can still be wrong where it's allowed to soften a verdict, which we accept only because we limit it to cases where a wrong answer is low-cost.

From signals to a verdict: the ordered cascade

How an ordered list of reviewed rules turns evidence into one verdict. Before the rules, a budget check: over budget returns HTTP 429, never a verdict. The same evidence — sender identity, link analysis and content label — is then read by every rule in turn. A rule reads the evidence and does not match; the next rule also does not match; the first rule that matches decides the verdict on its own and attaches its reason code, so a high-risk verdict always names the rule behind it. Every rule below it is never consulted for that message. A model's label is evidence a rule reads, never a verdict on its own, and the LLM second opinion runs inside one rule and only where a wrong yes costs little. Rule order is policy, so reordering is reviewed and tested like any other code change. The result is one verdict decided by exactly one rule, which the phone then parses against an allowlist.

The verdict comes from an ordered list of rules, where the first match decides and a high-risk rule always names itself (I3). Model labels and the second opinion are evidence that individual rules read, so no model decides a verdict on its own (I4).

We did consider a weighted score. A score tells you how bad something looks, but the cascade also tells you why, which is what users and analysts need, and policy changes become a matter of ordering reviewed, tested rules rather than tuning weights. A refusal never masquerades as a verdict either, so a caller over budget gets an HTTP 429 instead (I2).

The trade-off is that rule order is policy, and a new rule can quietly make a later one unreachable, so reordering gets the same scrutiny as code.

The last mile: verdicts on the phone

Plenty of writing about scam detection stops when the model answers. For us that's about halfway, because the phone still decides what the answer, or the lack of one, means for what a person can tap.

Every message on the phone is in exactly one state, one row per message. An SMS arrives and becomes pending — shown at once with links off, and notified once the scoring call returns or fails, with cleaned text. A strict contact match instead makes it a saved contact, which has links on and is notified, and is scored only if the user turns on personal chats. If there is no usable answer a retry sweep runs online only, oldest first, in batches, rescoring rows still pending. Answers pass through an allowlist parser accepting HIGH, LOW or UNKNOWN; anything else is no answer and the row stays pending. Low risk has links on, is notified and shows a category label. Unknown means the server says it cannot tell: links off, and final, so there is no retry storm on a sick service. High risk is hidden behind a placeholder with no notification, and a late verdict swaps alerts for a text-free warning. A failed check shows as Not checked when a deadline has passed, the message is gone or the sender is blocked. A first-run backfill puts past messages through the same parser in batches, marked read and without alerts.

Every message is in exactly one state. Unknown is reserved for the server to say it cannot classify something, and the phone treats it as final, because retrying would cause what the PR author called "a retry storm aimed at a service that is already down". Answers from the live call, the retry sweep and the first-run backfill all go through the same parser.

// Simplified for this post.
enum class Verdict { HIGH, LOW, UNKNOWN }

/** A verdict we recognise, or null for "no answer". Never a guess. */
fun parseVerdict(raw: String?): Verdict? =
    Verdict.entries.firstOrNull { it.name == raw?.trim()?.uppercase() }

/** Links are earned. Only a saved contact or a low-risk verdict turns them on. */
fun linksEnabled(verdict: Verdict?, savedContact: Boolean, hiddenChars: Boolean): Boolean =
    !hiddenChars && (savedContact || verdict == Verdict.LOW)

Both exist because of a whole class of bug. Earlier versions checked for the one bad answer and treated everything else as clean, so an answer the phone didn't recognise read as safe. The September audit flagged it, and since then a message has had to earn its links.

State In the conversation Links Notification
Pending Shown Off Once the scoring call returns or fails
Saved contact Shown On Yes
Low risk Shown, with its category label On Yes
Unknown Shown Off Yes
High risk Hidden: "Message hidden for your safety" Off None, and a late verdict replaces earlier ones with a warning that leaves out the text
Failed check Shown as "Not checked" Off Only the one from arrival

The rest of the phone-side design follows the same idea.

  • Hidden characters. A display sanitizer removes zero-width, direction-override and similar characters. Zero-width joiners stay because Devanagari and emoji need them, but their presence turns links off, and scoring and reporting always use the original text.
  • Notifications. Release 1.1.51 stopped holding back notifications for unscored messages, because OTPs never arrived when the network was down. A notification now waits only for the scoring call itself. A high-risk verdict means no notification, and anything else, including a failed call, notifies with cleaned text and a Copy OTP button. If a retry later finds a message high risk, the app clears that sender's notifications and posts a warning without the message text.
  • Retries. We never retry a POST in the network layer, since a timeout doesn't prove the request failed. A WorkManager sweep retries while the phone is online instead, with conditional writes that only touch rows still pending.
  • Contacts. Only a sender that could really be a phone number can match a saved contact, so a business-style sender ID can never borrow a friend's trust.
  • Consent. Every call carrying message data checks for consent and sign-in (I5), and without them messages are still delivered but stay pending.

Each rule lives in a small, pure policy object tested in CI. The verdict parser has 6 tests, the live-scoring callback 9, the link policy 8, the sanitizer 13 and the retry sweep 22, and that coverage is why we can say the app fails closed.

There are still gaps. An offline phone leaves a message pending, with its links off and no verdict. Messages from saved contacts are also not scored by default, so a scam sent from a friend's compromised number is trusted unless the user turns on scanning for personal chats.

Humans in the loop

Our analysts work through three queues in the admin portal, covering high-risk messages to verify, unknown headers to promote and registry renames to approve. People can also forward suspicious messages, screenshots and PDFs to our WhatsApp line, where an AI drafts a reply and an analyst edits and sends it, so no AI-written verdict reaches anyone there unchecked.

Design considerations

Pillar How it shows up
Security An explicit threat model, certificate pinning with a per-hop host guard, fenced LLM channels and bounded regular expressions
Privacy Consent gates (I5), the narrow OTP path described above, and a CI check that fails the build when a log call is handed message text, an OTP or a URL
Reliability Delivery first (I1), idempotent retries, a timeout on every upstream call, and refusals never mistaken for verdicts (I2)
Explainability One rule per verdict (I3), and an app that shows "Not checked" instead of staying silent
Cost An in-house model on CPUs doing most of the work, cached answers, usage budgets and no second-opinion or OTP calls for anonymous requests (I6)

Lessons learned

  • The classifier was the easy part. What the phone does with an answer, or without one, decides whether anyone is protected.
  • Parse values that cross a trust boundary against an allowlist, including your own server's answers, so anything you don't recognise means "no answer". Both fail-open bugs here checked for the one bad value and treated everything else as safe.
  • Keep anything an attacker can type away from the decision, and keep facts in channels they cannot write to.
  • Make promises structural, because a test can check an object's structure but nobody reliably remembers a flag.
  • Protect the data you trust most, since a sender registry is exactly what an attacker would want to poison.

Conclusion

Detecting a scam message is a classification problem, but protecting the person who receives it is a systems problem that depends on who sent the message, what a model may decide and what happens when an answer is late, wrong or missing. Most of it is unglamorous engineering, and little of it is specific to SMS. In upcoming posts we'll cover reading forwarded screenshots and PDFs without letting them instruct a model, and the registry pipeline.

If you're holding the phone

To spot a fake SMS, look at what it asks you to do. If a message pushes you to act within minutes, to pay, to call a number it gives you or to install something, treat it as a scam message until you've checked it some other way. Never share an OTP with anyone who asks for it, whoever they claim to be. If a link in EnfoldAI won't open, that's on purpose, because we haven't cleared it yet. If you've already lost money, call 1930, India's cyber-fraud helpline, or report it at cybercrime.gov.in as soon as you can.

What we leave out

We're happy to explain how our defences work, but not how to get around them. That's why this post leaves out thresholds, the rules themselves and their order, the exact conditions that trigger each check, model and vendor names, and hosting details. None of our protection depends on keeping the design secret, but there's no reason to hand out a tuning manual either.

Sources

Acknowledgments

This system is the work of EnfoldAI's Android, backend and data teams, and of the analysts who review what it flags. We especially thank our founding engineers, Krishna Agrawal and Gaurav Agarwal.

SecurityMachine learningAndroidBackend