RecentlyPostedJobs

Senior Software Development Engineer, Ads AI Core Infrastructure (ACI), Ads AI Core Infrastructure

Amazon · New York, New York, USA · United States

Employer
Amazon
Requisition id
10574802
First posted (employer ATS)
(19h ago)
First seen by this site
2026-10-09T17:52:59Z
Last verified live
2026-10-09T18:15:21Z
Source
Employer career portal (amazon)

Job description

At Amazon Ads, we're re-imagining the advertising landscape through advanced generative AI technologies and AI agents, revolutionizing how millions of customers discover products and engage with brands online. We are at the forefront of re-inventing advertising experiences, bridging human creativity with artificial intelligence to transform every aspect of the advertising lifecycle — from ad creation and optimization to performance analysis and customer insights. The Ads Agent Fabric team delivers the AI infrastructure platform that accelerates agentic application development from months to weeks. We build the agent-builder SDK with pre-integrated Ads features, developer tools that accelerate the develop, test, and model evaluation lifecycle, and a fully-managed stateful runtime with integrated access to Ads services, data, and foundational models. Our platform solves core production challenges: managing execution state across multi-step reasoning and multi-agent orchestration, enabling resumability through checkpointing and task queues for long-running workflows, and providing memory primitives that enable stateful experiences where agents remember and learn across sessions. You will own the architecture of Tier-1 services that handle millions of requests per day while holding millisecond latencies and strict SLAs. You will set the technical direction for distributed, service-oriented platforms powering the next generation of AI-driven advertising, and make the build-versus-buy and foundational technology choices that the team and its partner teams build on for years. Our systems and algorithms operate on massive datasets using distributed frameworks. In this role, you will design new and re-architect existing solutions for the fast-growing scale of agentic infrastructure — including stateful runtime services and pre-built platform components that agent-builder teams rely on to ship production agents. You will not merely go through the full software development cycle but more importantly drive appropriate technology choices for the business, lead the way for continuous innovation, and shape the future of AI-powered advertising at Amazon. We are a highly motivated, collaborative, and fun-loving team building at the forefront of AI agent infrastructure. We are entrepreneurial and have a bias for action with a broad mandate to experiment and innovate. This is an opportunity to make a significant impact on the future of AI agents in the Amazon Advertising business. A successful candidate will have the satisfaction of seeing their work power thousands of AI agents serving millions of advertisers globally, broaden their technical skills across distributed systems and ML infrastructure, and work in an environment that thrives on creativity, experimentation, and product innovation. Key job responsibilities Senior SDEs are technical leaders for their team and beyond. You take large, ambiguous problems that span multiple systems and teams and break them into designs other engineers can execute. You own the architecture of the team's most critical services and are accountable for their correctness, scalability, latency, and operational health over the long term. You drive the technology choices that matter - data and storage models, enforcement-path latency budgets, and the seams between the orchestration runtime, knowledge base, and governance layers and you document the trade-offs so decisions outlast any one project. You raise the quality bar across the team through design and code review, and you make hard calls on where to invest, what to simplify, and what to deprecate. You multiply the team's output: you mentor SDE-1s and SDE-2s, grow their design judgment, and set engineering standards that keep the platform maintainable as it grows. You identify systemic risks before they become incidents, drive root-cause fixes that eliminate whole classes of failure, and partner across Ads org boundaries to align contracts and roadmaps. You influence the team's technical roadmap, represent the team in cross-team and cross-org technical forums, and translate research from Applied Scientists into production-ready services. About the team Agent Fabric is Amazon Ads' AI infrastructure platform. We extend AWS AgentCore with advertising-specific capabilities and a managed, stateful runtime, cutting agent development from months to weeks. We own the platform end-to-end: the agent-builder SDK, the runtime that executes agents in production, and the knowledge and governance layers that keep it safe and accountable at scale. Our systems are in the critical path for GenAI projects across Amazon Advertising. Stateful Runtime & Orchestration: an orchestration framework that turns a single agent turn into durable, resumable work: async conversations that persist after a turn ends and notify on completion, scheduled and event-driven workflows, and multi-step DAG execution with fan-out, cancellation, and retry. Checkpointing, task queues, and short- and long-term memory let agents resume and learn across sessions. Knowledge Base: the RAG Knowledge Base that grounds advertiser answers across Sponsored Products, Sponsored Brands, DSP, reporting, and optimization. We run the KB agent, an OpenSearch vector index, and a daily ingestion pipeline spanning published and partner-curated content, with eval gates that enforce retrieval quality before every rebuild. Pre-Built Ads Components: delegated advertiser authorization, privacy-safe logging, policy guardrails, real-time data access, and context handoff between agents. Metering & Throttling: metering every invocation across model calls, tool calls, memory operations, and session time, then enforcing per-user and per-advertiser limits on a sub-millisecond path to protect the platform from abuse. Operational Excellence: parallelization for low latency, human-in-the-loop approvals, end-to-end tracing, security isolation, cost attribution, and versioned deploys with rollback.

Apply on Amazon’s site

More from Amazon

We are not Amazon. The hiring company owns this listing. Reposts of the same requisition id are not shown as new.

All new jobs · Companies we watch · How dates work · Report an error