WhatsApp's New Scam Alert Reads Your Messages Without Leaving Your Phone
WhatsApp is testing a feature that flags suspicious messages using an AI model that runs entirely on your device. Signal is also rolling out automatic checks to confirm nobody has secretly intercepted your conversations.

Key points
- WhatsApp began a limited beta test of its Scam Alert feature in mid-2025, flagging suspicious messages from unknown senders using a model that runs on the user's device.
- No message content is sent to WhatsApp or Meta; only two anonymised counts, how often warnings fire and what users do next, leave the phone.
- Every version of the AI model must be logged on an independent public record before WhatsApp can push it to devices, to prevent a targeted model being sent to a single person.
- Signal separately launched automatic key verification, using Cloudflare and Trail of Bits as independent auditors to confirm nobody has secretly inserted themselves into a private conversation.
- Both features are in early rollout and will continue to change based on feedback.
Two of the world's most-used encrypted messaging apps shipped privacy features this week aimed at the same basic problem: criminals using private channels to trick or spy on ordinary people.
What does WhatsApp's Scam Alert actually do?
It reads incoming messages on your phone, looking for patterns that match known scams, then shows you a private warning if something looks suspicious. Nothing leaves your device.
WhatsApp describes the tool as sitting "alongside" end-to-end encryption, meaning encryption (the system that scrambles messages so only sender and recipient can read them) stays intact. A small AI model, downloaded to your phone when you switch the feature on, scans messages from people not already in your contacts. It looks at things like conversational structure and word choices associated with fraud.
Only you see the warning. The sender has no idea. From there you can block the contact, report the message, dismiss the warning, or mark the chat as trusted so future alerts for that person are suppressed. If you mark a chat as trusted you can also, separately, choose to send the last five messages received to WhatsApp to help improve the model.
Can WhatsApp use this to spy on my chats?
The design makes that significantly harder than with traditional cloud-based filters, but the verification system is the real check on abuse.
Because no message text travels to WhatsApp's servers, the company built what it calls a confidential federated analytics pipeline, which is a system that collects only broad counts: how many warnings were shown across all users, and how many led to a block or a trust. Individual message content is never part of that.
More interestingly, every release of the AI model must be recorded on a third-party append-only transparency ledger (a public log that entries can be added to but never deleted or altered) before any device accepts it. Each release also comes with a cryptographic fingerprint, a set of SHA-256 hashes (unique digital signatures for each file), verified against keys held by Cloudflare rather than Meta. The goal is to stop WhatsApp quietly pushing a special version of the model to a specific person.
Users can inspect a log of which messages were scanned and which model version made each call under Account, Request Info, then Scam Alert Activity.
| Feature | Where processing happens | Data sent externally | Independent oversight |
|---|---|---|---|
| Scam Alert model | On your device | Anonymised counts only | Cloudflare holds signing keys |
| Model distribution | On your device | None | Third-party transparency ledger |
| Federated analytics | On your device | Aggregate counts | Threat model published by WhatsApp |
| Signal key verification | On your device | None | Cloudflare, Trail of Bits |
What is Signal doing differently?
Signal launched automatic key verification, a system that checks, without any action from you, that nobody has secretly placed themselves between you and the person you are messaging.
End-to-end encrypted apps use public keys (think of a public key as a padlock only you hold the key to) to secure conversations. If an attacker broke into Signal's systems and swapped your padlock for their own, they could read your messages. Signal's new key transparency system, as it is called, logs every key registration and change in a public cryptographically verifiable record, meaning the record itself proves it has not been tampered with.
Cloudflare and Trail of Bits act as independent auditors of that log. All user identifiers in it are scrambled, so the auditors confirm the log's integrity without ever seeing anyone's phone number or username in plain text.
The feature is on by default. Users who prefer to do manual safety number checks, which involve comparing a string of numbers with a contact in person or over another channel, can turn automatic verification off under Privacy, Advanced, then Automatic Key Verification.
First reported by SecurityWeek, both features are described by their companies as early previews subject to change.
What should you do now?
For WhatsApp users: the beta is limited, so Scam Alert may not be available to you yet. When it arrives, enabling it costs nothing and the privacy trade-off is narrow. If you receive a scam warning, treat it as a prompt to slow down, not a guarantee the message is fraudulent.
For Signal users: automatic key verification is already rolling out. A green "Encryption verified" badge on a contact's profile means the check passed. No action is required unless you want to opt out.



