Software Architect, GPU System
ByteDance · San Jose
- Employer
- ByteDance
- Requisition id
- A241145
- First posted (employer ATS)
- (6d ago)
- First seen by this site
- 2026-10-06T18:18:43Z
- Last verified live
- 2026-10-06T19:47:55Z
- Source
- Employer career portal (bytedance)
Job description
About the Team
We are a systems software team building the foundational software for large-scale compute platforms. We work at the hardware/software boundary across the Linux kernel, accelerators, storage, firmware, and platform validation. We value rigorous engineering, clear interfaces, measurable performance and reliability, and upstream collaboration where appropriate. The team partners closely with hardware, architecture, product, validation, and production engineering groups to move new capabilities from design through dependable deployment.
About the Role
You will lead the technical strategy and cross-functional delivery of the low-level GPU software stack for data-center computing. You will connect accelerator architecture, firmware, kernel drivers, runtimes, communication, telemetry, and fleet reliability into a coherent roadmap. This is a hands-on individual-contributor role: you will make architecture decisions, guide engineers, review critical implementations, and resolve system-level issues without assuming hiring or performance-management responsibilities.
Responsibilities
- Define the architecture and multi-year technical roadmap for GPU system software across driver, runtime, firmware-interface, management, observability, and reliability layers.
- Set priorities and technical standards for accelerator enablement, balancing near-term product commitments with compatibility, performance, serviceability, and long-term maintenance.
- Own end-to-end integration from architecture and pre-silicon planning through bring-up, qualification, general availability, fleet monitoring, and sustained operation.
- Lead cross-functional execution among silicon, firmware, kernel, compiler, library, machine-learning framework, server, network, storage, validation, and production teams.
- Resolve ambiguous system-level tradeoffs involving APIs, resource management, memory and interconnect topology, telemetry, recovery, security, and workload performance.
- Remain technically involved through prototypes, critical-path code and design reviews, performance analysis, and leadership during high-severity failure investigations.
- Define measurable release and reliability criteria, including qualification coverage, regression thresholds, fault containment, automated repair, and fleet-health indicators.
- Mentor engineers, raise the quality of architecture and debugging practices, and communicate technical decisions to both specialist and executive audiences.
Minimum Qualifications
- Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience.
- 5+ years of hands-on experience in GPU, accelerator, kernel, firmware, runtime, or data-center system software, including technical leadership on complex hardware/software programs.
- Demonstrated success setting technical direction and leading a major hardware/software program from concept through production.
- Deep understanding of accelerator architecture, memory hierarchy, interconnects, operating systems, device management, and production reliability.
- Experience coordinating dependencies and technical decisions across multiple engineering teams and organizational boundaries.
- Strong coding, design-review, performance-analysis, and system-debugging capability, with evidence of continued hands-on contribution.
Preferred Qualifications
- Experience with a major GPU computing stack and its kernel driver, runtime, libraries, tooling, and fleet-management interfaces.
- Experience defining platform APIs or hardware/software contracts across multiple accelerator generations.
- Knowledge of distributed accelerator workloads, collective communication, topology-aware placement, and large-scale serviceability.
- Track record mentoring senior engineers and building alignment where priorities, ownership, or technical evidence initially conflict.
More from ByteDance
- Senior Software Engineer, Storage Systems 6d ago
- Senior Software Engineer, GPU Systems 6d ago
- System Software Architect, Linux Kernel and Operating System 6d ago
- Software Engineer (Merchant Product) - Payment Product and Solution - Global Payment 7d ago
- Software Engineer (SRE - Platform Services), Infrastructure Engineering 14d ago
- Full-stack Software Engineer - BytePlus 15d ago
- Senior Software Development Engineer, Storage Engine 27d ago
- Software Development Engineer, Storage Engine 27d ago
We are not ByteDance. The hiring company owns this listing. Reposts of the same requisition id are not shown as new.
All new jobs · Companies we watch · How dates work · Report an error