- Lab
-
Libraries: If you want this lab, consider one of these libraries.
- Cloud
Troubleshoot AWS Workload Health Using Observability Signals
In this AWS lab, you investigate a workload health issue using observability signals in the AWS Console, without changing any application code. Starting from a CloudWatch alarm on a content-delivery service backed by Amazon S3, you compare metrics and logs to decide whether the alarm reflects real user impact or noise, trace the elevated error rate to its likely source, and recommend an operational next step. You work the way an on-call engineer does: read the alarm and the reported symptom, line up a metric against a log, separate the requests that are actually failing users from incidental noise, and turn the evidence into a defensible finding. By the end, you can take a raw alert and state clearly what is happening, what the evidence shows, and what to do next.
Lab Info
Table of Contents
-
Challenge
Establish the condition under investigation
- Review the alarm and the reported symptom to state what condition requires investigation.
- Open the workload health view and note the time window the signal falls in and how broad it is.
-
Challenge
Compare the metric and the change history to characterize the condition
- Examine the client-error metric against total request volume to gauge the magnitude and timing of the signal.
- Examine the recent change history for the service and note the management event that lines up with the onset of errors.
- Reconcile the metric and the change history into a single characterization of the condition.
-
Challenge
Trace the symptom to its source and separate impact from noise
- Confirm the failing pattern directly by attempting representative reads and observing which are refused and which succeed.
- Establish that the refused reads are explained by the change identified in the history, as the likely source.
- Distinguish the reads that represent real user impact from those that show the issue is scoped rather than a full outage.
-
Challenge
Summarize the finding and the next operational action
- State the likely cause and the evidence that supports it.
- Identify the recommended next operational action.
About the author
Real skill practice before real-world application
Hands-on Labs are real environments created by industry experts to help you learn. These environments help you gain knowledge and experience, practice without compromising your system, test without risk, destroy without fear, and let you learn from your mistakes. Hands-on Labs: practice your skills before delivering in the real world.
Learn by doing
Engage hands-on with the tools and technologies you’re learning. You pick the skill, we provide the credentials and environment.
Follow your guide
All labs have detailed instructions and objectives, guiding you through the learning process and ensuring you understand every step.
Turn time into mastery
On average, you retain 75% more of your learning if you take time to practice. Hands-on labs set you up for success to make those skills stick.