Sphinx 2026h2 Crash and Misconfiguration

773 views
Skip to first unread message

Rick Roos

unread,
Jul 18, 2026, 12:10:20 AM (6 days ago) Jul 18
to Certificate Transparency Policy
At approximately 1:30 AM on July 17th UTC time the Sphinx 2026h2 log saw an increase in traffic for get-entries that caused load on its CTile cache layer. The CTile cache layer did not have the correct HPA configuration to allow the pods to automatically scale with the increased load. This caused stress on the CTile pods which eventually caused them to crash, which in turn caused more traffic to route to the Trillian front-end pods, also causing them to crash.

While trying to update configuration to increase the CTile pod count, an incorrect value was included in the routing for the get-entries endpoint, which caused any get-entries requests to be routed to a CTile cluster that was configured to return entries from a test log instead of the Sphinx 2026h2.

Ultimately this means between 2:11 AM and 2:34 AM on July 17th UTC time, any API calls to get-entries that were in the range of 0–4850122 (inclusive) would have received entries from the wrong log. Any requests for entries higher than that range would have received an HTTP 400.

As of now, the log is running normally with the correct routing configuration, and the CTile pods are appropriately scaled. The log is still accepting entries.

Thanks,
Rick

Andrew Ayer

unread,
Jul 20, 2026, 9:00:06 AM (3 days ago) Jul 20
to Rick Roos, Certificate Transparency Policy
Hi Rick,

Since Friday, the get-entries endpoint has been so severely rate-limited that I can only get 17 entries per second from the log. Meanwhile, the log has continued to accept entries and issue SCTs at a rate above 120 entries per second.

Regards,
Andrew
> --
> You received this message because you are subscribed to the Google
> Groups "Certificate Transparency Policy" group. To unsubscribe from
> this group and stop receiving emails from it, send an email to
> ct-policy+...@chromium.org. To view this discussion visit
> https://groups.google.com/a/chromium.org/d/msgid/ct-policy/CACrB8xkw9BwcS25MM7h3p6xS9H%3DTXDKwpo6O_kKVXUwfJ_mF8w%40mail.gmail.com.

Andrew Ayer

unread,
Jul 21, 2026, 2:56:00 PM (2 days ago) Jul 21
to ct...@digicert.com, Certificate Transparency Policy
+ ct...@digicert.com

For over 4 days this log has grown at 7x the rate at which it is possible to retrieve entries. Its 24h availability has dropped below 90% as measured by Google[1].

Regards,
Andrew

[1] https://www.gstatic.com/ct/compliance/endpoint_uptime_24h.csv
> https://groups.google.com/a/chromium.org/d/msgid/ct-policy/20260720090000.63f080dc513be1ccfa74f304%40andrewayer.name.

Colin Stubbs

unread,
Jul 22, 2026, 12:07:05 AM (yesterday) Jul 22
to Andrew Ayer, ct...@digicert.com, Certificate Transparency Policy

Seconding this as it's still the case at the moment.

Growth is far exceeding anything that can be achieved with the rate limit being imposed at the moment.

Available requests per second seems to at most ~0.75. Any faster than that and I see 429's.

Sphinx2026h1 and it is not hitting the same rate limit as I have a system actively catching up on that still. Sphinx2027h1/h2 seems fine as well.

This seems to be a very specific rate limit for https://sphinx.ct.digicert.com/2026h2/ct/v1/get-entries that has been forgotten.

-Colin
> To view this discussion visit https://groups.google.com/a/chromium.org/d/msgid/ct-policy/20260721145554.658bde64d07b3dcbb34563c5%40andrewayer.name.
>

Andrew Ayer

unread,
Jul 22, 2026, 3:17:53 PM (yesterday) Jul 22
to Certificate Transparency Policy, ct...@digicert.com
This is a 5 day long incident which DigiCert has not even acknowledged.

If this log cannot serve entries to monitors within the MMD, then it is no longer serving a purpose to the ecosystem and needs to be Retired.

Regards,
Andrew

On Tue, 21 Jul 2026 14:55:54 -0400
> https://groups.google.com/a/chromium.org/d/msgid/ct-policy/20260721145554.658bde64d07b3dcbb34563c5%40andrewayer.name.

Ben Cartwright-Cox

unread,
Jul 22, 2026, 4:26:35 PM (yesterday) Jul 22
to Certificate Transparency Policy, Andrew Ayer, ct...@digicert.com
If it helps as a data point, the CT monitors I run have also been
hopelessly behind since BST (UTC+1) midnight on the 18th on
sphinx.ct.digicert.com/2026h2 and wyvern.ct.digicert.com/2026h2 has
been spotty
> To view this discussion visit https://groups.google.com/a/chromium.org/d/msgid/ct-policy/20260722151748.c9f0ac86ddb53446e623c5d5%40andrewayer.name.

Jeremy Rowley

unread,
Jul 22, 2026, 4:42:35 PM (yesterday) Jul 22
to Ben Cartwright-Cox, Certificate Transparency Policy, Andrew Ayer, ct...@digicert.com
Hi all, 

Sorry about the delay. The normal CT ops team was unavailable, hence the delayed response. We updated the CTtile scaling this morning. You should see faster through-put. We're investigating whether this increase is enough to handle the monitor requests coming in and will adjust accordingly, providing further updates as needed.  

Jeremy

Ben Cartwright-Cox

unread,
Jul 22, 2026, 4:48:51 PM (yesterday) Jul 22
to Certificate Transparency Policy, Jeremy Rowley, Andrew Ayer, ct...@digicert.com
From my end I've seen no performance improvement

Google's uptime currently reads sphinx 2026h2 at 89.2% availability...

> https://sphinx.ct.digicert.com/2026h2/,get-entries,89.2857

My understanding of the CT programme is that alone should trigger
sphinx 2026h2 to be retired

Jeremy Rowley

unread,
Jul 22, 2026, 4:54:25 PM (yesterday) Jul 22
to Ben Cartwright-Cox, Certificate Transparency Policy, Andrew Ayer, ct...@digicert.com
Other logs have fallen below 90% and survived (Sectigo has a log recently that did this) so availability is not the instant removal that it used to be.  The log is still operational  and none of our alarms are triggering as it is failing for monitors, not CAs or on our internal systems. We're investigating that and will adjust the log's operation as needed.  Of course, if Google wants to retire the log then that's fine as well. However, it would be good to know  if it will be retired right away so we don't waste time troubleshooting and fine-tuning the logs scale.  

We had a huge spike of certs being logged starting the 16th, which seems to be the cause of the issues.  

Ben Cartwright-Cox

unread,
Jul 22, 2026, 5:02:14 PM (yesterday) Jul 22
to Jeremy Rowley, Certificate Transparency Policy, Andrew Ayer, ct...@digicert.com
(Tossing the fellow CT operator hat and putting on the "I run a side
company that tracks CT logs as well)

I'm trying to say this is politely I can, but have you tried to
monitoring/following your own logs? It is extremely difficult to stay in
sync with the latest entries

Jeremy Rowley

unread,
Jul 22, 2026, 5:05:52 PM (yesterday) Jul 22
to Ben Cartwright-Cox, Certificate Transparency Policy, Andrew Ayer, ct...@digicert.com
Yes - we have an internal monitor pulling certs right now. It is falling behind so we can definitely see what you're seeing when it comes to adding certs faster than we can pull them. I'm hoping we can make some more adjustments so the monitors can catch up with the increased rate of logging. If we can up the tile scaling enough that you can pull faster than the new certs are being logged, then everything should go back to status quo. I'm not on that engineering team but have asked for an update on when we'll see a change to improve the log performance for monitors. 

Filippo Valsorda

unread,
Jul 22, 2026, 5:18:13 PM (yesterday) Jul 22
to Jeremy Rowley, Ben Cox, Certificate Transparency Policy, Andrew Ayer, ct...@digicert.com
2026-07-22 22:54 GMT+02:00 Jeremy Rowley <rowl...@gmail.com>:
Other logs have fallen below 90% and survived (Sectigo has a log recently that did this) so availability is not the instant removal that it used to be.

This is a decidedly concerning response.

Log operators commit to a policy which requires 99% availability. (Note that the 24h uptime is informational. The policy measures a 90d rolling window, in which this log currently has 99.1662% availability.) The fact that logs were allowed to violate that policy in the past was an unfortunate necessity when logs were critically few, it's not a permanent license to disregard it. As a community, we should strive to improve the health of the ecosystem, not take opportunities to ratchet it down.

The log is still operational  and none of our alarms are triggering as it is failing for monitors, not CAs or on our internal systems.

Claiming that the log is operational is even more puzzling: monitors are the main customers of CT logs. If monitors report being unable to keep up with the log, the log is unavailable. The reason MMDs exist is to bound the time between certificates being issued and certificates reaching monitors. If monitors are more than 24h behind, that is an ongoing outage.

We're investigating that and will adjust the log's operation as needed.  Of course, if Google wants to retire the log then that's fine as well. However, it would be good to know  if it will be retired right away so we don't waste time troubleshooting and fine-tuning the logs scale.  

This feels misplaced. Google made the requirements very clear. If you are unable or unwilling to meet them, you should retire the log, not let Google do it for you, and on top of that demand they deliver that decision urgently for your convenience.

I know this might be strident coming from a fellow CT log operator, but the whole reason Geomys runs a log is to demonstrate that it is possible to run a compliant log with a limited budget, and help fellow operators do so. If well-funded operators argue against the importance of the bare-minimum availability bar, that kind of defeats the point.

Jeremy Rowley

unread,
Jul 22, 2026, 5:31:14 PM (yesterday) Jul 22
to Filippo Valsorda, Ben Cox, Certificate Transparency Policy, Andrew Ayer, ct...@digicert.com
Other logs have fallen below 90% and survived (Sectigo has a log recently that did this) so availability is not the instant removal that it used to be.

This is a decidedly concerning response.

Log operators commit to a policy which requires 99% availability. (Note that the 24h uptime is informational. The policy measures a 90d rolling window, in which this log currently has 99.1662% availability.) The fact that logs were allowed to violate that policy in the past was an unfortunate necessity when logs were critically few, it's not a permanent license to disregard it. As a community, we should strive to improve the health of the ecosystem, not take opportunities to ratchet it down.


No, it's not a concerning response. The response is pointing out that historically this 99% has not triggered an immediate removal of the log.  If this was a requirement that mandated immediate removal, then we would turn down the log now. The history with logs means we aren't sure if this log is not viable or not based solely on availability to monitors. 
 
The log is still operational  and none of our alarms are triggering as it is failing for monitors, not CAs or on our internal systems.

Claiming that the log is operational is even more puzzling: monitors are the main customers of CT logs. If monitors report being unable to keep up with the log, the log is unavailable. The reason MMDs exist is to bound the time between certificates being issued and certificates reaching monitors. If monitors are more than 24h behind, that is an ongoing outage.

That is not correct either. MMD is the mean time to merge, not the mean time to pull.  It could be considered an outage I suppose, but you can't redefine MMD.

We're investigating that and will adjust the log's operation as needed.  Of course, if Google wants to retire the log then that's fine as well. However, it would be good to know  if it will be retired right away so we don't waste time troubleshooting and fine-tuning the logs scale.  

This feels misplaced. Google made the requirements very clear. If you are unable or unwilling to meet them, you should retire the log, not let Google do it for you, and on top of that demand they deliver that decision urgently for your convenience.

We can retire the log. Retiring logs is easy. However, again based on precedent, not keeping up with monitors is not a shutdown worthy event by default. MMD has not been missed.  
 
I know this might be strident coming from a fellow CT log operator, but the whole reason Geomys runs a log is to demonstrate that it is possible to run a compliant log with a limited budget, and help fellow operators do so. If well-funded operators argue against the importance of the bare-minimum availability bar, that kind of defeats the point.

Not really relevant to whether the log has failed is it? We could shut this one down and spin up a new shard, but I don't think its necessary given the policy and how previous drops in Google-measured availability have been handled.

Andrew Ayer

unread,
Jul 22, 2026, 5:36:54 PM (yesterday) Jul 22
to Ben Cartwright-Cox, 'Ben Cartwright-Cox' via Certificate Transparency Policy, Jeremy Rowley, ct...@digicert.com
On Wed, 22 Jul 2026 22:02:07 +0100
"'Ben Cartwright-Cox' via Certificate Transparency Policy"
<ct-p...@chromium.org> wrote:

> I'm trying to say this is politely I can, but have you tried to
> monitoring/following your own logs?

I'll add that there are many more signals that log operators can use to determine that there is a problem with their log:

- Chrome's uptime CSV feed, and Geomys' service which makes it easy to set up alerts on it: https://uptime.geomys.org/ct/

- The error feeds published by SSLMate at https://sslmate.com/resources/certspotter_stats as well as the CSV which includes backlogs https://feeds.sslmate.com/ct_logs.csv

- crt.sh's backlogs, which can be accessed through PostgreSQL

Like Ben I've observed no improvement in download rate today.

All of the points raised by Chrome to DigiCert last year <https://groups.google.com/a/chromium.org/g/ct-policy/c/aR6gKzCANVs/m/7qCOGa8XCAAJ> could apply equally to this incident: log operators should proactively monitor their logs for problems, multi-day delays in response are bad, and a log's write API should be proactively disabled while remediating an incident.

Regards,
Andrew

Jeremy Rowley

unread,
Jul 22, 2026, 5:39:50 PM (yesterday) Jul 22
to Andrew Ayer, Ben Cartwright-Cox, 'Ben Cartwright-Cox' via Certificate Transparency Policy, ct...@digicert.com
We are monitoring our logs for problems, and we saw an improvement from the rate limit changes.  I can't comment on the lack of response (as I'm not on the team that normally responds to these), but I have requested that we disable the log's write API already.  I've asked for confirmation when its complete so I can share that here.

Filippo Valsorda

unread,
Jul 22, 2026, 5:45:50 PM (yesterday) Jul 22
to Jeremy Rowley, Ben Cox, Certificate Transparency Policy, Andrew Ayer, ct...@digicert.com
2026-07-22 23:30 GMT+02:00 Jeremy Rowley <rowl...@gmail.com>:
Other logs have fallen below 90% and survived (Sectigo has a log recently that did this) so availability is not the instant removal that it used to be.

This is a decidedly concerning response.

Log operators commit to a policy which requires 99% availability. (Note that the 24h uptime is informational. The policy measures a 90d rolling window, in which this log currently has 99.1662% availability.) The fact that logs were allowed to violate that policy in the past was an unfortunate necessity when logs were critically few, it's not a permanent license to disregard it. As a community, we should strive to improve the health of the ecosystem, not take opportunities to ratchet it down.


No, it's not a concerning response. The response is pointing out that historically this 99% has not triggered an immediate removal of the log.  If this was a requirement that mandated immediate removal, then we would turn down the log now. The history with logs means we aren't sure if this log is not viable or not based solely on availability to monitors. 
 

The log is still operational  and none of our alarms are triggering as it is failing for monitors, not CAs or on our internal systems.

Claiming that the log is operational is even more puzzling: monitors are the main customers of CT logs. If monitors report being unable to keep up with the log, the log is unavailable. The reason MMDs exist is to bound the time between certificates being issued and certificates reaching monitors. If monitors are more than 24h behind, that is an ongoing outage.

That is not correct either. MMD is the mean time to merge, not the mean time to pull.  It could be considered an outage I suppose, but you can't redefine MMD.

The MMD is the Maximum Merge Delay. I am not redefining it, I am explaining that the reason it exists is to bound the delay between a certificate obtaining an SCT and monitors being able to access it. The availability requirement serves the same purpose, and so does the Rate Limiting section of the Chrome Certificate Transparency Log Policy.

The log did not violate its MMD, but monitors are delayed, which means that the system is not working. That is an incident, and it is concerning and perplexing not to see it taken seriously.

Jeremy Rowley

unread,
Jul 22, 2026, 5:48:48 PM (yesterday) Jul 22
to Filippo Valsorda, Ben Cox, Certificate Transparency Policy, Andrew Ayer, ct...@digicert.com
It's being taken seriously for sure. Just later than we would have liked. I've got the team looking at it now and asked them to shut off new inclusions until the root cause (and fix) is found.  After we figure out how to ensure everyone is caught up, then we can take a look at what failed in the communication back to the community. As mentioned, we could just call this shard dead if that's the preference. However, I'm not sure what is considered dead with the change in how uptime is treated.

Cynthia Revström

unread,
Jul 22, 2026, 5:50:51 PM (yesterday) Jul 22
to Jeremy Rowley, Andrew Ayer, Certificate Transparency Policy, ct...@digicert.com
Hi Jeremy,

Like others I am very concerned with how DigiCert seemingly isn't
treating this like an on-going incident, or at least wasn't initially.

In the thread that Andrew linked in response to the previous incident
you clearly said the following in response to Chrome's disappointment
in your delayed response:
> Yes - looking back through the communication, we clearly failed to address the incident and communication from Google in a timely manner. Our SLAs for responding to incidents (regardless of source) are being updated with an expectation of communication within 1 hour and (where possible) resolution within 24 hours.

Clearly that SLA didn't seem to matter.

I would very much like to hear why despite being criticized for this
just last year you failed on the same point again.
If you aren't the correct person to answer this question, could you
get the correct person to come here and respond to it?

-Cynthia (not a CT log operator or monitor, just an interested party)
> --
> You received this message because you are subscribed to the Google Groups "Certificate Transparency Policy" group.
> To unsubscribe from this group and stop receiving emails from it, send an email to ct-policy+...@chromium.org.
> To view this discussion visit https://groups.google.com/a/chromium.org/d/msgid/ct-policy/CAFK%3DoS-N%2BgASTbi%3D23Qt4J-Pcot%2BeaSa%2B%3DTLapw5%3D_CWbhq_kA%40mail.gmail.com.

Jeremy Rowley

unread,
Jul 22, 2026, 5:55:45 PM (yesterday) Jul 22
to Cynthia Revström, Andrew Ayer, Certificate Transparency Policy, ct...@digicert.com
You're right that it wasn't being treated as an incident initially, and your concern is fair. As far as timelines go, the issue was escalated to both myself and the exec team today.  We've made it an urgent priority and are treating it as an incident now. 

> "I would very much like to hear why despite being criticized for this
just last year you failed on the same point again.
If you aren't the correct person to answer this question, could you
get the correct person to come here and respond to it?"

Yeah - I'm not the right person to answer that. We'll have someone address that directly from the CT team. Right now, I want to make sure that I'm keeping you all informed and pushing internally to get it fixed. I will have someone from the CT team provide an RCA of what happened and why we didn't communicate back right away. (I can't give more details myself because I don't know what happened yet.)

Joe DeBlasio

unread,
Jul 22, 2026, 6:29:13 PM (24 hours ago) Jul 22
to Jeremy Rowley, Cynthia Revström, Andrew Ayer, Certificate Transparency Policy, ct...@digicert.com

Hi folks,


We had not chimed in on this thread until now because our compliance monitoring infrastructure still has sphinx2026h1 above 99% 90-day availability, albeit barely. However, I wanted to clarify our view on the importance of log availability, and how we form our policy responses.


I would direct attention to Chrome's CT Log Policy section on rate limiting. In particular, it reads in part, "Log data availability is of paramount importance. All rate limits must be set to ensure that all well-behaved clients can reliably retrieve log entries at a rate greater than the growth rate of the log."


The policy also notes that "removal of the log from Chrome’s log list may not be the best outcome for Chrome, Chrome’s users, or the CT ecosystem." As a result, dropping below 99% availability does not automatically result in the retirement of a log by our policy.


In the CT ecosystem of a year ago, we were deeply concerned that the removal of several logs risked the viability of the overall ecosystem. We had 25% fewer log operators then, with the average log offering a much lower quality of service than logs today. Ultimately, our policy decisions, now and then, consider both the ecosystem's health, and all information available from the incident, including how the operator responds to the incident. I'd direct attention to the corresponding policy section on incident detection and response.


Thanks for working to address the issue, Jeremy. I hope we can look forward to a prompt resolution, and assuming sphinx2026h2's availability does indeed drop below 99%, a full postmortem once the technical issues have been addressed. 


Best,

Joe, on behalf of the Chrome CT Team



Jeremy Rowley

unread,
Jul 22, 2026, 6:35:25 PM (24 hours ago) Jul 22
to Jeremy Rowley, Cynthia Revström, Andrew Ayer, Certificate Transparency Policy, Certificate Transparency Operations
FYI - the log was moved to read only at 4:23 mountain pending our investigation and fix in whatever is causing logging to exceed the rate monitors can pull. 



From: Jeremy Rowley <rowl...@gmail.com>
Sent: Wednesday, 22 July 2026 15:55:25
To: Cynthia Revström <m...@cynthia.re>
Cc: Andrew Ayer <ag...@andrewayer.name>; Certificate Transparency Policy <ct-p...@chromium.org>; Certificate Transparency Operations <ct...@digicert.com>
Subject: Re: [ct-policy] Sphinx 2026h2 Crash and Misconfiguration
 

Jeremy Rowley

unread,
Jul 22, 2026, 9:12:28 PM (21 hours ago) Jul 22
to Jeremy Rowley, Cynthia Revström, Andrew Ayer, Certificate Transparency Policy, Certificate Transparency Operations

Status right now

As mentioned, Sphinx 2026h2 write access is off as of 4:23 PM Mountain today, July 22nd. We will keep it off until until we have a fix for the read throttling and confirmation that monitors have caught up. We are looking at Wyvern to see if that log has the same issue. 

What we've done so far is temporarilty remove the read throttling on Sphinx get-entries.  That produced a short window where pull rates improved. After about half an hour, the log came under enough read load that it started to fail, and we had to reimpose a limit. The limit we put in place is higher than the one Andrew and Colin have been hitting. Andrew, Colin, Ben - if you could check again and let me know if you're seeing improvement then that would help. We are watching the numbers overnight ourselves and will adjust as needed, but I thought your observations would be useful in determining the read-write balance. 

What caused this

Certificate volume on Sphinx 2026h2 spiked starting July 16th. Write volume was manageable. The read side was not withget-entries traffic running at roughly 50 percent higher volume than write traffic. On top of that, each read request pulls a batch of entries compared to the one certificate per write, making the read data close to 100x the write data. The rate limit on get-entries was not sized for that increase, which is why the monitors have been hitting the limits over the last few days. 

Why we did not catch this sooner

This is the part Cynthia asked about directly, so I want to answer it straight rather than fold it into the technical explanation. Note this is my initial analysis and not the fulll write-up of what happened. Details will likley change as the investigation is on-going.

Rick's July 17th post described the CTile misconfiguration as resolved. We treated the elevated read latency that followed as a residual symptom of that incident and spike in logging rather than a separate, ongoing one. Monitors falling behind the logging rate is its own issue and should have been separated out from the July 16th post. 

Separately, our CT mailing list coverage failed at the same time. We keep three people on rotation monitoring the list. Two were on approved personal leave and one was pulled into a personal emergency with no advance notice. The people covering for them did not realize this thread needed urgent handling, so it sat as a routine item instead of an incident.  This was because none of our internal alerts were firing as the system was not considered "down". That is a gap in our on-call handoff process. 

As mentioned, our internal alerts did not fire because they are built around service availability, not around read throughput relative to monitors. We have an alert for monitors but only that the read-access is working, not that the read-access speed has dropped below write-access speed. The log was considered up, even if access was too slow for monitors to keep pace. 

We're still working on optimizing access but wanted to give you all an update on where we are at and the preliminar findings. 

Jeremy

Colin Stubbs

unread,
1:59 AM (16 hours ago) 1:59 AM
to Jeremy Rowley, Certificate Transparency Policy, Certificate Transparency Operations
Thanks Jeremy,

I can confirm I see this has improved now as of the last 3 hours or so, as of about 2026/07/23 00:22:21 UTC.

It seems more like a 3 rps limit now?

Rather than what seemed like a 0.3 rps limit previously.

-Colin

Cynthia Revström

unread,
7:28 AM (11 hours ago) 7:28 AM
to Jeremy Rowley, Jeremy Rowley, Certificate Transparency Policy, Certificate Transparency Operations
Thank you for the update, it mostly answers my initial question.
It also raises some questions that I will hold-off on until the
incident is resolved and the write-up is out.

-Cynthia

Jeremy Rowley

unread,
12:16 PM (6 hours ago) 12:16 PM
to Cynthia Revström, Jeremy Rowley, Certificate Transparency Policy, Certificate Transparency Operations
We’ve turned write access back on for Sphinx. We’re working on the full incident report now.
Reply all
Reply to author
Forward
0 new messages