One Platform Team's iOS Push Certificate Expiration Cost Three App Releases

Jul 17, 2026 By Lucas Mendes

In early 2025, Team Alpha at a fintech startup supporting a consumer app with millions of users found themselves in a familiar but painful situation: a certificate had expired. Not just any certificate—the Apple Push Notification service (APNs) certificate. The team had been rotating other certificates quarterly, but the push cert was on a different renewal cycle. When it expired, Apple's App Store Connect refused to accept new builds. Three consecutive release trains were skipped. The cost: roughly 6–8 weeks of delayed feature rollouts and an estimated 40–60 person-days of emergency re-signing, re-approval, and lost productivity.

One Expired Certificate, Three Missed Deadlines

The platform team responsible for iOS code signing and provisioning had a well-oiled process for rotating development and distribution certificates. They used Fastlane match to store and sync signing assets across machines. Every quarter, someone ran a renewal script. But push notification certificates were not part of that rotation. They lived in a separate section of the Apple Developer Portal, with a different expiry date that no one had documented.

When the push cert expired, the first sign of trouble came from a failed build. Fastlane match could sign the app binary, but App Store Connect rejected the upload because the push entitlement referenced an invalid certificate. The team spent a day diagnosing the issue, only to realize the cert had expired two days earlier. Renewal required generating a new CSR, uploading it to the Developer Portal, and downloading the new cert—then updating the server-side push provider. The entire process took several hours, but the real bottleneck was App Review.

Each resubmission required a new build, which meant going through App Review again. The first missed release was a minor feature update. The second was a critical bug fix. The third included a new subscription tier that had been in development for months. By the time the cert was renewed and all three builds were approved, roughly six weeks had passed. The team estimated that engineering time spent on the incident—including diagnosis, renewal, re-signing, re-submission, and post-mortem—totaled around 50 person-days.

The incident was not unique. A 2019 outage at a major ride-sharing app locked out push notifications for millions of users due to a similar expiry. In that case, the company's backend failed to fall back gracefully, and users received no ride updates for hours. The push cert had been rotated years earlier by a single engineer who had since left the company. No one else knew where the cert was stored or when it expired.

The Hidden Chain of Dependencies in Mobile CI/CD

Mobile app development involves multiple layers of signing and identity. For iOS, the chain includes development certificates, distribution certificates, provisioning profiles, and push notification certificates. Each has its own expiry date, renewal process, and dependency on Apple's Developer Portal. Android uses Google Play Signing for app distribution and Firebase Cloud Messaging (FCM) keys for push, which also expire. Cross-platform teams using Flutter or React Native must manage two or three separate certificate lifecycles per app.

The problem is not that certificates expire—it's that expiration is often invisible until a build fails. Apple does not provide bulk expiration alerts for Developer Portal certificates. The portal shows expiry dates in a table, but there is no API to query all certificates and get a consolidated view. Many teams rely on calendar reminders or manual audits. Some use Fastlane's cert or match commands to list certificates, but those tools only show what is stored locally, not what is active on the portal. Developers can set up email notifications for certain events (e.g., certificate renewal), but these are not reliable. Many teams have resorted to building their own monitoring, either by scraping the portal or by using third-party services that check certificate expiry daily.

Automated CI/CD pipelines often lack checks for certificate expiration. A typical pipeline might run unit tests, build the app, sign it, and upload it to App Store Connect. If the signing certificate is valid, the build succeeds. But if a push cert expires between builds, the next upload will fail. The failure is not caught in the build step; it surfaces during validation on Apple's servers, which can take several minutes. The pipeline then reports a cryptic error about entitlements, leaving developers to trace the issue.

Android faces a similar issue with FCM keys. Google's Firebase Console provides an expiry date, but the key is often shared across multiple services. When it expires, push notifications silently stop working. The app does not crash, and the developer may not notice until users complain. Some teams have implemented monitoring that sends a test push notification every day and alerts if delivery fails. But such monitoring is not standard practice.

How One Platform Team Discovered the Gap

The team in question discovered the gap through a series of alerts. The first alert came from Fastlane match after a build failed because the push certificate was not found in the match repository. The team had been rotating certificates quarterly but had never added the push cert to match. It was stored in a shared folder on a developer's machine, with no backup. The second alert came from the CI system, which reported that the upload to App Store Connect had been rejected. The third alert came from the product manager, who asked why the release had not gone out.

The post-mortem revealed that the team had missed similar expirations in two other services: a VoIP certificate for a communication feature and a Wallet certificate for digital passes. Both had been renewed manually after the fact, but no one had tracked the root cause. The inventory of all certificates was kept in a spreadsheet, which had not been updated in six months. The team did not have a single source of truth for certificate metadata—only scattered entries in spreadsheets, email threads, and notes.

The incident prompted the team to audit all certificates across the organization. They found that several internal services used self-signed certificates that had expired years ago. Those services still worked because the clients did not validate the expiry date. But the push certificate was validated by Apple's servers, so expiration was enforced. The team realized that certificate management was not just a developer concern—it was an infrastructure reliability issue that required a systematic approach.

One engineer on the team had previously worked at a larger company where certificate lifecycle was managed through an internal certificate authority (CA). That company issued short-lived certificates that were automatically renewed by a service. For push certificates, they used Apple's Developer Portal but wrote a script that checked the expiry date daily and sent a Slack alert 30 days before expiration. The team adopted a similar approach after the incident.

The Real Cost: Developer Hours and Lost Revenue

The direct cost of the expired push certificate was measured in developer hours. The team estimated that the incident consumed roughly 50 person-days across six engineers. That included time spent diagnosing the issue, renewing the certificate, re-signing three builds, re-submitting to App Review, waiting for approval (2–4 days per release), and writing the post-mortem. The opportunity cost was higher: three sprints of feature work sat unreleased, including a new subscription tier that could have generated revenue.

In-app purchases were delayed by roughly six weeks. Assuming the app had a modest conversion rate and average revenue per user, the lost revenue could have been substantial. The team did not publish exact figures, but based on typical industry benchmarks, a delay of this magnitude could cost a mid-sized app tens of thousands of dollars in lost subscription revenue alone. The team also incurred intangible costs: team morale dipped as launch dates slipped publicly. The product manager had to communicate delays to stakeholders, and the engineering team felt the pressure of emergency work.

The incident also highlighted the cost of manual processes. Renewing a push certificate involves generating a new CSR, uploading it to the Developer Portal, downloading the new cert, updating the server-side push provider (e.g., AWS SNS, Firebase Cloud Functions), and updating the app's provisioning profile. Each step is manual and error-prone. The team had to coordinate with the backend team to update the push provider, which required a separate deployment. The entire process took several hours, even though the actual renewal took minutes.

Some teams have attempted to quantify the cost of certificate expiration more broadly. A 2023 survey by the mobile DevOps vendor Bitrise found that 40% of respondents had experienced a production incident related to certificate expiry in the past year. The median cost per incident was estimated at $10,000–$50,000, including engineering time and lost revenue. For larger apps with millions of users, the cost could be higher. The survey also found that only 25% of teams had automated certificate renewal for push services.

Industry Patterns: Certificate Expiry Is a Known Risk

Certificate expiry is a well-documented risk in the industry. In 2019, a major ride-sharing app experienced a push notification outage when its APNs certificate expired. The company's backend had been using the same certificate for years, and no one had set a reminder. The outage lasted several hours and affected millions of users. The company later implemented a monitoring system that checked certificate expiry daily and rotated certificates automatically using Fastlane match.

Several open-source tools exist to help manage certificate lifecycles. Fastlane's match tool stores signing assets in a git repository and can automatically renew certificates when they expire. However, match only covers certificates that are part of the Apple Developer Program, not push certificates stored on the server side. For push certificates, tools like cert-manager (for Kubernetes) or custom scripts can be used. Adoption is uneven, with many teams still relying on manual audits.

Larger organizations often use an internal certificate authority (CA) for internal certificates, but push certificates must come from Apple. Some companies have negotiated with Apple to get longer-lived certificates (e.g., three years instead of one), but this is rare. The standard validity period for APNs certificates is one year. For FCM keys, the validity is typically two years. Teams must plan for renewal cycles that align with their release cadence.

Automating Certificate Lifecycle Management

After the incident, the team implemented several automation measures. First, they added the push certificate to Fastlane match, which stored the certificate in a git repository and allowed automatic renewal. They configured match to run a weekly check that compared the expiry date of stored certificates against a threshold. If any certificate was within 30 days of expiry, match would generate a new one and update the repository.

Second, they added an expiration check to their CI pipeline. Before building the app, the pipeline would query the push certificate's expiry date using a script that parsed the certificate file. If the certificate was within 14 days of expiry, the pipeline would fail with a clear message. This prevented builds from being submitted with soon-to-expire certificates. The team also added a scheduled job that checked all certificates in the organization and sent a Slack alert 30, 14, and 7 days before expiry.

Third, they integrated certificate monitoring with their incident response system. Critical certificates—those that could block a release—were monitored by PagerDuty. If a certificate expired, an alert would be sent to the on-call engineer. The team also created a runbook for emergency re-signing, which included steps for renewing the certificate, updating the server-side push provider, and re-submitting the app. The runbook was tested quarterly in a staging environment.

Finally, they established a shared vault for certificate secrets using a tool like HashiCorp Vault. The vault stored the private keys and certificates with least-privilege access. Only the platform team had write access, while other teams could read the public certificates. The vault also stored metadata about each certificate, including the expiry date, owner, and dependencies. This replaced the spreadsheet and provided a single source of truth.

Lessons for Platform Teams of Any Size

The incident offers several lessons for platform teams. First, treat certificates as infrastructure, not configuration. Certificates have lifecycles that require monitoring, renewal, and testing. They should be managed with the same rigor as servers, databases, or network devices. A spreadsheet is not a reliable inventory tool. Use a dedicated certificate manager or a secrets vault with expiry tracking.

Second, inventory all certificates and their dependencies in a single source of truth. Include not just signing certificates but also push certificates, VoIP certificates, Wallet certificates, and any other certificate used by the app or its backend services. Document which service uses each certificate, the renewal process, and the team responsible. Review the inventory quarterly alongside feature planning.

Third, test the renewal process in staging before production expiry. Many teams discover that renewal requires steps they had forgotten, such as updating the push provider or regenerating a CSR. A dry run in staging can uncover these issues without blocking a release. The team in question now performs a quarterly renewal drill for all certificates, rotating them even if they haven't expired, to ensure the process works.

Fourth, build a runbook for emergency re-signing. Even with automation, things can go wrong. A runbook should include step-by-step instructions for renewing the certificate, updating the server, and re-submitting the app. It should also include contact information for the Apple Developer support team and the push provider. Test the runbook at least once a year.

Finally, review lifecycle policies quarterly alongside feature planning. Certificate expiry dates should be visible in the team's planning tool. If a certificate is due to expire during a critical release window, the team should plan to renew it early. Some teams have adopted a policy of renewing all certificates at the start of each quarter, regardless of their expiry date. This reduces the risk of missing a renewal and simplifies the process.

What would your team do if a certificate expired tomorrow? Would you catch it before the build failed, or would you be scrambling to re-sign and re-submit? The answer depends on whether you've invested in the automation and processes that make certificate expiry a non-event. For Team Alpha, the three missed releases were a painful but necessary wake-up call. For other teams, the lesson is clear: treat certificates as critical infrastructure, or risk paying the price in delayed releases and lost trust.

Recommend Posts
Tech

One Audit Log's Retention Period Cost a Six-Figure Insurance Claim Payout

By Yusuke Tanaka/Jul 17, 2026

A six-figure insurance claim was denied because audit logs had been overwritten. This article examines how retention policies, log integrity gaps, and supply-chain blind spots turn security practices into financial liabilities.
Tech

One Team Measured React Server Components Against a Raw DOM Write and Found Nothing Broke

By Lucas Mendes/Jul 17, 2026

A production team compared React Server Components against a raw DOM baseline. Two weeks, 1.2 million sessions, and no regressions. Here's what they learned.
Tech

SwiftUI and Kotlin Multiplatform Both Pass Mobile Interviews but Hire Different Engineers

By Lucas Mendes/Jul 17, 2026

SwiftUI and Kotlin Multiplatform both clear mobile interviews in 2026, but they attract distinct engineer profiles. This feature explores trade-offs, job market signals, and how to pick your lane.
Tech

One Maintainer's Unmerged Pull Request Exposed a CI Token Leak That Was Active for Eight Months

By Deepa Iyer/Jul 17, 2026

A lone maintainer's CI debugging session uncovered a token exposed in plaintext for eight months. The unmerged PR reveals systemic gaps in supply-chain security.
Tech

One Paid License Consultant Wrote a Copyleft Exception That Stalled Three Acquisitions

By Sara Park/Jul 17, 2026

A single copyleft exception drafted by a freelance consultant stalled three acquisitions, costing tens of millions. How one bad clause became a poison pill.
Tech

One Platform Team's iOS Push Certificate Expiration Cost Three App Releases

By Lucas Mendes/Jul 17, 2026

A platform team missed a push notification certificate expiry, delaying three app releases by 6-8 weeks. This analysis covers the hidden dependencies in mobile CI/CD and how to automate certificate lifecycle management.
Tech

Flutter's Widget Tree vs SwiftUI's View Body Two Teams Paid for Both

By Deepa Iyer/Jul 17, 2026

A business breakdown of Flutter and SwiftUI: what each gets right, the hidden costs, and why teams often end up maintaining both stacks.
Tech

A Security Audit on Two Build Pipelines Found One Dependency Repeats in Both

By Deepa Iyer/Jul 17, 2026

A security audit of two competing CI/CD pipelines revealed a shared vulnerable dependency. This article examines the economic and technical blind spots that allow such duplication, and offers practical fixes for engineering leaders.
Tech

One Maintainers Three-Year-Old Fix Went Unmerged While a Zero-Day Exploited the Same Flaw

By Deepa Iyer/Jul 17, 2026

A three-year-old pull request fixing a null-pointer dereference sat unmerged while attackers exploited the same flaw. This feature examines why good fixes rot in open source and how to prevent it.
Tech

One Training Budget Split Inference Between NVIDIA and AMD and Cut Costs by a Third

By Sara Park/Jul 17, 2026

Splitting inference across NVIDIA and AMD GPUs can cut costs by a third. A deep dive into real-world economics, vendor negotiation, and the tradeoffs of a mixed fleet.
Tech

React Server Components and HTMX Both Offer Less JS But One Team Quit

By Lucas Mendes/Jul 17, 2026

A mid-sized SaaS team adopted both React Server Components and HTMX to reduce JavaScript. Half the engineers quit within six months. Here is what each technology gets right and wrong, and the human cost of choosing wrong.
Tech

Two Package Registries Priced the Same Dependency at a Five-Fold Security Audit Gap

By Sara Park/Jul 17, 2026

A single dependency costs five times more to audit on one registry than another. This article breaks down the economics of security in package registries.
Tech

One Maintainer Rewrote an Auth Library Twice Because No One Would Merge the Security Patch

By Sara Park/Jul 17, 2026

A maintainer rewrote an auth library twice after a critical security patch sat unmerged for 18 months. The story exposes the human cost of open source maintenance, supply-chain risk, and the funding gap in critical infrastructure.
Tech

SwiftUI and Jetpack Compose Share One Syntax But Two Team Cultures

By Deepa Iyer/Jul 17, 2026

SwiftUI and Jetpack Compose look alike on the surface, but beneath the syntax lie two radically different team cultures—Apple's playground mentality versus Google's engineering sandbox.
Tech

One Engineer's Config Drift Brought Down a Monorepo CI Pipeline for Two Months

By Deepa Iyer/Jul 17, 2026

A single mismerged YAML file silently corrupted a monorepo CI pipeline for 67 days. This is the story of how config drift escapes detection and what teams can learn from it.
Tech

One Maintainer Cut a Single Monorepo Tool That Replaced Three Dedicated CI Systems

By Yusuke Tanaka/Jul 17, 2026

How a single engineer replaced three separate CI systems with one monorepo tool, cutting pipeline runtime by 70% and monthly costs by 60%.
Tech

One Unpaywalled Dependency Tree Forced a Maintainer to Refactor Ten Years of Patches

By Deepa Iyer/Jul 17, 2026

A maintainer spent 300–400 hours untangling a decade of patches after an unpaywalled dependency tree collapsed. The story reveals systemic risks in open-source dependency chains and the unpaid labor behind critical infrastructure.
Tech

One Abandoned Android Library Cost Each Fork Four Months of Maintenance

By Yusuke Tanaka/Jul 17, 2026

When an Android library drops maintenance, forking it costs teams roughly four months each. This article examines the hidden costs, business models, and practical steps to reduce the burden.
Tech

One Inference Engineer's GPU Swarm Saved a Week per Pipeline Run

By Deepa Iyer/Jul 17, 2026

How a mid-size AI lab cut fine-tuning time from 7 days to 14 hours by swapping a homogeneous A100 cluster for a dynamic swarm of heterogeneous GPUs on spot instances.
Tech

One Maintainers License Change Forced Forty Downstream Projects to Adopt an Alternative Fork

By Yusuke Tanaka/Jul 17, 2026

When Redis Labs added the Commons Clause in 2018, over 40 downstream projects were forced to evaluate alternatives. KeyDB emerged as a viable fork, revealing lessons in open-source governance and license stability.