RecentlyPostedJobs

Software Architect, GPU System

ByteDance · San Jose

Employer
ByteDance
Requisition id
A241145
First posted (employer ATS)
(6d ago)
First seen by this site
2026-10-06T18:18:43Z
Last verified live
2026-10-06T19:47:55Z
Source
Employer career portal (bytedance)

Job description

About the Team We are a systems software team building the foundational software for large-scale compute platforms. We work at the hardware/software boundary across the Linux kernel, accelerators, storage, firmware, and platform validation. We value rigorous engineering, clear interfaces, measurable performance and reliability, and upstream collaboration where appropriate. The team partners closely with hardware, architecture, product, validation, and production engineering groups to move new capabilities from design through dependable deployment. About the Role You will lead the technical strategy and cross-functional delivery of the low-level GPU software stack for data-center computing. You will connect accelerator architecture, firmware, kernel drivers, runtimes, communication, telemetry, and fleet reliability into a coherent roadmap. This is a hands-on individual-contributor role: you will make architecture decisions, guide engineers, review critical implementations, and resolve system-level issues without assuming hiring or performance-management responsibilities. Responsibilities - Define the architecture and multi-year technical roadmap for GPU system software across driver, runtime, firmware-interface, management, observability, and reliability layers. - Set priorities and technical standards for accelerator enablement, balancing near-term product commitments with compatibility, performance, serviceability, and long-term maintenance. - Own end-to-end integration from architecture and pre-silicon planning through bring-up, qualification, general availability, fleet monitoring, and sustained operation. - Lead cross-functional execution among silicon, firmware, kernel, compiler, library, machine-learning framework, server, network, storage, validation, and production teams. - Resolve ambiguous system-level tradeoffs involving APIs, resource management, memory and interconnect topology, telemetry, recovery, security, and workload performance. - Remain technically involved through prototypes, critical-path code and design reviews, performance analysis, and leadership during high-severity failure investigations. - Define measurable release and reliability criteria, including qualification coverage, regression thresholds, fault containment, automated repair, and fleet-health indicators. - Mentor engineers, raise the quality of architecture and debugging practices, and communicate technical decisions to both specialist and executive audiences. Minimum Qualifications - Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience. - 5+ years of hands-on experience in GPU, accelerator, kernel, firmware, runtime, or data-center system software, including technical leadership on complex hardware/software programs. - Demonstrated success setting technical direction and leading a major hardware/software program from concept through production. - Deep understanding of accelerator architecture, memory hierarchy, interconnects, operating systems, device management, and production reliability. - Experience coordinating dependencies and technical decisions across multiple engineering teams and organizational boundaries. - Strong coding, design-review, performance-analysis, and system-debugging capability, with evidence of continued hands-on contribution. Preferred Qualifications - Experience with a major GPU computing stack and its kernel driver, runtime, libraries, tooling, and fleet-management interfaces. - Experience defining platform APIs or hardware/software contracts across multiple accelerator generations. - Knowledge of distributed accelerator workloads, collective communication, topology-aware placement, and large-scale serviceability. - Track record mentoring senior engineers and building alignment where priorities, ownership, or technical evidence initially conflict.

Apply on ByteDance’s site

More from ByteDance

We are not ByteDance. The hiring company owns this listing. Reposts of the same requisition id are not shown as new.

All new jobs · Companies we watch · How dates work · Report an error