# OPS06-BP04 - Automate testing and rollback

Best practice: OPS06-BP04
Pillar: Operational Excellence
Source: https://wellarchitected.cloudvisor.eu/docs/operational-excellence/ops06-bp04.html

## Implementation Guidance

"Automate testing and rollback" should be implemented through codified workflows, not ad hoc manual steps. Prioritize idempotent automation, failure handling, and rollback controls so teams can operate safely at scale.

For the question "How do you mitigate deployment risks?", define measurable outcomes, assign owners, and review execution regularly. Integrate this practice into delivery and operations processes so improvements persist as workloads and requirements evolve.

### Key Steps

1. **Design automation boundaries**:
   - Identify which parts of "Automate testing and rollback" should be fully automated
   - Define pre-checks, post-checks, and approval controls
   - Specify rollback behavior and exception handling requirements

2. **Implement and integrate workflows**:
   - Codify automation in pipelines, runbooks, or event-driven handlers
   - Add telemetry, alerting, and audit trails for each automated action
   - Validate idempotency and safe re-execution under failure conditions

3. **Harden and continuously improve**:
   - Run failure simulations to validate automation behavior
   - Track error rates, execution time, and manual fallback frequency
   - Refine logic and controls based on incident and operations feedback

## Risk / Impact

**Level of risk if not implemented**: High

**Impact**: If this best practice is missing, teams are more likely to experience preventable incidents, delayed recovery, and inconsistent change outcomes. Control gaps and weak visibility can increase customer impact during high-pressure events.

**Benefits of implementation**:
- Reduced operational risk through repeatable controls
- Faster detection and response during incidents
- Stronger auditability and decision traceability

## AWS Services to Consider

<div class="aws-service">
  <div class="aws-service-content">
    <h4>AWS CodeDeploy</h4>
    <p>Deploys application updates with strategies such as canary and linear rollout.</p>
  </div>
</div>

<div class="aws-service">
  <div class="aws-service-content">
    <h4>AWS CodePipeline</h4>
    <p>Automates release workflows with quality gates and controlled promotions.</p>
  </div>
</div>

<div class="aws-service">
  <div class="aws-service-content">
    <h4>Amazon CloudWatch</h4>
    <p>Collects metrics, logs, and alarms that support operational insight and performance management.</p>
  </div>
</div>

<div class="aws-service">
  <div class="aws-service-content">
    <h4>Elastic Load Balancing</h4>
    <p>Distributes traffic across healthy targets for better availability and response time.</p>
  </div>
</div>

## Related Resources

<div class="related-resources">
  <h2>Related Resources</h2>
  <ul>
    <li><a href="https://docs.aws.amazon.com/wellarchitected/latest/framework/ops-06.html">OPS06: How do you mitigate deployment risks?</a></li>
    <li><a href="https://docs.aws.amazon.com/wellarchitected/latest/operational-excellence-pillar/welcome.html">AWS Well-Architected Framework - Operational Excellence Pillar</a></li>
    <li><a href="https://docs.aws.amazon.com/wellarchitected/latest/operational-excellence-pillar/design-principles.html">Operational Excellence Design Principles</a></li>
  </ul>
</div>
