Rollback, recovery and integrity: UN R156 and ISO 24089
What happens when an update fails — and how R156 and ISO 24089 expect you to handle it
The moment an update fails
Every OTA and workshop update programme is defined by what it does when things go wrong, not by the clean case. A rollback and recovery design is the set of behaviours that keeps a vehicle safe and usable when an update is interrupted, corrupted, or found to be faulty after installation. UN R156 does not treat this as an edge case to be handled later; safe failure is one of the properties the regulation expects the update process to have, and it is one of the things an assessor looks for.
The failures are mundane and inevitable. Power drops mid-flash. Connectivity is lost part-way through a download. A package arrives corrupted. An installed update turns out to misbehave in the field. A serious SUMS treats each of these as a designed branch of the workflow, with a defined outcome and a record, rather than an exception nobody planned for.
Integrity comes first
Rollback and recovery only make sense on top of integrity. If the vehicle cannot tell whether the software it holds is authentic and unaltered, it cannot know whether it needs to roll back, nor whether the image it would roll back to is trustworthy. So the first requirement is that the vehicle verifies the integrity and authenticity of an update before installing it, and that it retains a known-good image or recovery path it can trust. Integrity verification is the gate; rollback and recovery are what happens when that gate — or the installation after it — does not pass. This is why the topic sits inside the wider secure OTA update workflow rather than beside it.
Rollback versus recovery
The two words are often used loosely, but they describe different mechanisms, and a robust design usually needs both.
| Aspect | Rollback | Recovery |
|---|---|---|
| What it does | Returns to the previous, verified software | Returns to a defined safe operating state |
| When it applies | A new image installed but is faulty or fails verification | A flash is interrupted and the target is in an incomplete state |
| Precondition | A retained known-good previous image | A minimal trusted boot/recovery capability |
| Risk if absent | Fleet stuck on a bad version | Control unit left undefined or bricked |
| Record required | Outcome and version reverted to | Outcome and state recovered to |
Rollback answers "the new software is bad — go back to the old one." Recovery answers "the update did not complete — get to a safe state and, ideally, retry." An update that can roll back but cannot recover from an interrupted flash still risks leaving a control unit unusable; an update that can recover but cannot roll back a bad-but-installed image still risks stranding the fleet on a defective version. Design for both.
The safe-state is defined in advance
The recurring theme in R156 is that safe behaviour is designed and verified before release, not improvised at the moment of failure. The safe-state a vehicle enters when an update fails must be defined ahead of time — continue on the previous verified software, or restrict to a limited but safe mode until recovery completes — and it must be chosen so that no unsafe or undefined behaviour results. Pre-conditions on installation exist for the same reason: an update should not begin unless the vehicle can complete it safely or fail safely if it does not.
Where ISO 24089 does the engineering
ISO 24089:2023, "Road vehicles — Software update engineering", is where the abstract requirement becomes an engineering discipline. R156 says the update must fail safely and keep records; ISO 24089 provides the organisational and project processes to design, implement and verify that behaviour — including how failure modes are identified, how the safe-state and recovery paths are specified, and how they are tested before a campaign goes live. We cover the standard as a whole in ISO 24089 explained; for rollback and recovery specifically, it is the difference between asserting that your update fails safely and being able to show, with test evidence, that it does.
Records close the loop
Whatever happens — clean install, rollback, or recovery — the outcome has to be recorded. For each affected vehicle the SUMS should be able to say what was attempted, whether it succeeded, and if not, what state the vehicle ended in. These records matter twice: they are the evidence an assessor opens, and they are what lets a manufacturer answer a field-safety question about a specific vehicle months after the campaign. An update programme that rolls back correctly but keeps no record of having done so has solved the engineering problem and left the compliance problem open.
The failure that finds you at audit
The common gap is not the absence of a rollback mechanism — most modern update stacks have one. It is the absence of verification and records around it. The safe-state is asserted but never tested against a real interrupted flash. Recovery works on the bench but was never exercised on a representative failure. The outcomes are handled in code but not captured in the SUMS. Each of these passes a demo and fails an assessment, because the assessor asks not "does it roll back?" but "show me the evidence it rolled back safely on a real failure, and the record that proves it."
The AutoSifu view
We design rollback, recovery and integrity as one compliance-and-engineering problem: mapped to R156 and ISO 24089, built into the update architecture, and evidenced for CoC and VTA on a single route. With CIRT working alongside us, the approval body sees the failure-handling design and its test evidence while they can still be shaped, so what the assessor eventually opens is the record set that was planned from the start rather than assembled under pressure.
Questions
- What does UN R156 require on rollback?
- UN R156 requires that a software update process protect the vehicle and its user when an update does not complete successfully. In practice that means an update must fail safely: the vehicle must return to a known-good state rather than being left in an undefined or unsafe condition. Whether that is achieved by rolling back to the previous software or by recovering to a safe operating state, the outcome must be defined in advance and recorded.
- How does ISO 24089 support UN R156?
- ISO 24089:2023 is the software update engineering standard that provides the organisational and project-level processes UN R156 relies on. R156 states what must be true — safe failure, integrity, records — and ISO 24089 supplies the engineering practices that make it repeatable, including how update failure and recovery are designed and verified. Using it is not legally mandatory, but it is the practical basis for most SUMS evidence.
- What is a safe-state on update failure?
- A safe-state is a defined condition the vehicle enters when an update fails, chosen so that no unsafe or undefined behaviour results. It might mean continuing to run the previous, verified software, or restricting to a limited but safe operating mode until recovery completes. The key is that the state is defined and verified beforehand, not left to chance at the moment a flash is interrupted.
