
The worst failure in site maintenance is the one that does not look like a failure. The site loads, monitoring is green, the contact form submits and shows its thank-you message. Meanwhile the order notifications sit in the recipient's spam folder, and you learn about it two weeks later when the client asks why their CRM is empty.
The cause is almost always the same: somebody touched the DNS records that authenticate the domain's email. Here is how that breaks, why delivery monitoring does not catch it, and what routine prevents you hearing about it from the client first.
Receiving mail servers do not take a message's word for who sent it. Every inbound email is checked against three mechanisms, each with a distinct job.
SPF (a TXT record starting with v=spf1) answers "who is allowed to send
mail on behalf of this domain" — a list of servers and services.
DKIM (a TXT record at selector._domainkey.domain) is a cryptographic
signature. The sending server signs with a private key, the recipient verifies
with the public key from DNS. The signature proves the message was not altered
in transit.
DMARC (a TXT record at _dmarc.domain) ties the two together and tells the
recipient what to do with a message that fails: nothing (p=none), place it in
spam (p=quarantine), or reject it (p=reject).
One subtlety causes most of the confusion: SPF and DKIM can both pass while DMARC still fails. DMARC requires the domain that passed authentication to align with the domain in the From header. Sending "from your domain" through a service that signs with its own domain does not satisfy that. This is the source of nearly every "but I configured everything" conversation.
Per RFC 7208 a domain must have exactly one record beginning with v=spf1.
With two, the result is permerror — and that does not mean "one of them will
work", it means SPF does not pass at all.
The way it happens is mundane. The domain already has SPF for corporate mail. A
marketing tool is added, its setup guide asks for a record, and whoever does it
adds a second record instead of appending an include: to the existing one.
Both records look correct in isolation. Email breaks completely.
This is the nastiest variant, because the record stays syntactically valid.
The standard permits at most ten mechanisms that require a DNS query:
include, a, mx, ptr, exists, redirect. The limit is counted
recursively — if your include:_spf.example.com itself contains three more
include statements, you have spent four lookups, not one.
It looks like this:
v=spf1 include:_spf.google.com include:sendgrid.net include:spf.mandrillapp.com
include:_spf.crm-tool.com include:mail.hosting.com ~all
Five include statements in the record. But _spf.google.com contains three
more inside it, and sendgrid.net two. That is eleven against a limit of ten,
and SPF now fails for every sender, including the one that worked for years.
ip4: and ip6: mechanisms do not count — there is no DNS query in them.
So the way back under the limit is to replace some include statements with the
concrete address ranges, where the service publishes them.
+allv=spf1 +all means "anyone may send mail as this domain". It shows up as
debugging residue: someone set it to "just make it work" and never removed it.
It is a standing invitation to spoof the domain.
Strictly speaking this is not a breakage — domains ran without DMARC for decades. But since 2024 Gmail and Yahoo require it from bulk senders, and without a policy anyone can send mail as your domain with recipients having no way to tell.
Also worth understanding the policy difference. p=none means "send me reports
but deliver everything anyway". The record exists, DMARC is formally configured,
and it stops no spoofing whatsoever.
A record reading v=DKIM1; p= technically exists but contains no public key.
Per RFC 6376 an empty p= means the key is revoked or not published — the
selector is visible, but no signature will verify.
We ran into this while testing our own implementation: example.com returns
v=DKIM1; p= for any selector you ask for. The first version of our check
dutifully "found" all fourteen common selectors at once.
The intuitive way to monitor email is to send a message and confirm it arrives. That is exactly what we do elsewhere: a check submits a form carrying a unique token and waits for a message with that token in a control mailbox.
But broken SPF does not lose mail. It sends it to spam. The message arrives, the token is found, the delivery check reports success — and it is not lying, the message really did arrive. In a folder nobody reads.
The conclusion: end-to-end delivery checking and DNS record checking are two different checks, and neither substitutes for the other. The first answers "did it arrive", the second "will it arrive in the inbox".
Check email records when taking a site on retainer, after any DNS work, and then periodically — once a day is plenty, since records change by hand and rarely.
MX present. No MX records means the domain cannot receive mail at all. Obvious, yet routinely true of domains where mail is "about to be set up".
Exactly one SPF record. A quick look:
dig +short TXT example.com | grep "v=spf1"
Two lines back is an outage, even if both look correct.
DNS lookup count in SPF. It must be counted with nested include statements
expanded. Doing that by hand is tedious, which is precisely why this breakage
survives for months: nobody recounts the limit after adding another service.
A terminator. The record should end in -all (reject firmly) or ~all
(mark softly). Exception: if the record uses redirect=, the terminator comes
from the target record and its absence is not a fault.
DMARC present, and which policy. dig +short TXT _dmarc.example.com. If the
policy is p=none, log it as a task, not as a pass.
DKIM selectors. Here there is an irreducible difficulty: the selector is
listed nowhere in DNS. From the outside you can only guess common names
(default, mail, google, selector1, selector2, k1, s1) or take it
from the mail provider's settings. If you know the selector, check that one, and
treat a missing record as a fault.
We added a dedicated check type, "Email records (SPF/DMARC/DKIM)". It sends no mail and needs no mailbox access: it reads the domain's DNS and verifies everything listed above.
The check reports a failure when the failure is objective: no MX, no SPF, more
than one SPF record, the DNS lookup limit exceeded, +all present, DMARC
missing, an explicitly configured DKIM selector not found. The lookup limit is
counted with nested include statements expanded — that is, the way the
recipient's mail server will count it.
A fingerprint of the records — how many MX, how many of the ten lookups are spent, which DMARC policy, which selectors were found — is shown on a healthy check too. That is deliberate: when a contractor edits SPF, the change becomes visible in the check history even when the new record did not break the limit.
A useful side effect: the very first run against our own domain revealed that it had no DMARC record. The check began by finding a gap in its authors' setup.
permerror, not "one of them wins".Pingvera watches whether an online business actually works — uptime, checkout, orders, domain, SSL and server — and alerts you in Telegram, email or a webhook before a customer has to tell you.
Start free