Software Development Engineer, ML Acceleration, Trainium AI Systems, Annapurna Labs
Amazon · Austin, Texas, USA · United States
- Employer
- Amazon
- Requisition id
- 10570092
- First posted (employer ATS)
- (7h ago)
- First seen by this site
- 2026-10-06T06:17:27Z
- Last verified live
- 2026-10-06T06:47:31Z
- Source
- Employer career portal (amazon)
Job description
Annapurna Labs designs the silicon behind AWS machine learning acceleration. Trainium and Inferentia servers train and serve the largest models our customers run, and before any of that hardware carries customer traffic it has to prove it works. MLA Vetting owns that gate. We build the diagnostic tests that every accelerator server runs before it becomes sellable, and we decide which failures block a host, which route to a technician for repair, and which are noise.
That gate only works if we can see through it. We are hiring a Software Development Engineer II to build the data and analytics systems that tell us what our tests are actually doing across the fleet. You will build the pipelines that ingest diagnostic results from tens of thousands of servers, the dashboards that make test behavior legible, and the metrics and alarms that tell us a test has started failing on a hardware generation before it costs us a week of capacity. When the team decides whether a diagnostic has enough evidence to move from observation into blocking production, your data answers that question.
This is a software engineering role, not a reporting role. You will write and operate production code, own the alarms and metrics that page us, and build detection for trends nobody is watching yet. You will work directly with hardware, firmware, provisioning, and data center operations teams who act on what you build.
You are fluent in Python and SQL and have built production data pipelines at scale. Experience with Amazon Redshift, Apache Spark, workflow orchestration, and dashboarding in Grafana or QuickSight maps directly to what you will work on here. Experience with telemetry or diagnostic data from a large server fleet, or with defining metrics and anomaly detection on operational time-series data, will have you productive faster. Familiarity with server hardware or data center operations is a plus, not a requirement — we will teach you the hardware.
Key job responsibilities
Design, build, and operate production data pipelines that ingest diagnostic, telemetry, and repair-ticket data from the Trainium and Inferentia fleet into a warehouse other teams query with confidence.
Build and own the metrics, alarms, and anomaly detection that surface test regressions, failure-rate shifts, and new failure signatures across hardware generations without a human going looking for them.
Build dashboards and visualizations that make fleet and test health legible to engineers, hardware partners, and leadership, covering failure rates, failure-signature breakdowns, first pass yield, and repair latency.
Analyze large-scale fleet data to find root cause behind failure trends, and separate genuine hardware faults from software defects and test noise — a distinction that decides whether a failure reaches a technician or an engineer.
Define the evidence standard that gates operational decisions, including whether a diagnostic has soaked long enough and cleanly enough in the fleet to move from observation mode into blocking production.
Improve data quality and pipeline reliability so downstream consumers trust the numbers without re-deriving them.
Write clear analyses and design documents for technical and non-technical readers, including leadership.
A day in the life
You start by checking an alarm that fired overnight: a diagnostic's failure rate climbed on one server generation. You query the failure signatures, find the increase concentrates in one component and one data center, and hand the breakdown to the hardware team by mid-morning. The rest of the day goes to a pipeline you are building to join repair outcomes against diagnostic results, so the team can measure how often a repair actually fixes the fault it was dispatched for. You close the day reviewing a teammate's code review and answering an operations partner who needs a new cut of yield data.
About the team
MLA Vetting sits between hardware manufacturing and customer-ready capacity. We own the diagnostic test suites and the triage framework that turn a raw hardware failure into a specific, actionable repair, and we own the fleet-scale evidence that drives hardware and firmware improvements upstream. Our work is measured in how fast a server reaches sellable and how often it gets there on the first try.
More from Amazon
- Software Development Engineer II, EU INTech Partner Growth Experience 7h ago
- Software Development Engineer, Amazon Leo for Government 7h ago
- Principal Software Engineer, Amazon Traffic Engineering 7h ago
- Software Engineer II, MLA Automation & Deployments 7h ago
- Software Development Engineer II, EU INTech Partner Growth Experience 7h ago
- Software Development Engineer II, Int'l TechnologyEURegionalfund 7h ago
- Software Engineer-AI/ML, Inference Team - AWS Neuron 7h ago
- Software Development Engineer 1d ago
We are not Amazon. The hiring company owns this listing. Reposts of the same requisition id are not shown as new.
All new jobs · Companies we watch · How dates work · Report an error