LPU Chip Architecture Engineer

Il y a 3 jours

Brussels, Brussels-Capital, Belgique Canaan Inc. Temps plein
Responsibilities Lead the overall architecture definition of LPU chips based on the static dataflow architecture, including computing array design and on-chip SRAM storage hierarchy planning. Solve core pain points of high latency and frequent data movement in large model inference scenarios. Collaborate with the compiler team in the early stage to define hardware microarchitecture and implement hardware-software co-design. Align scheduling logic in advance to avoid industry pain points caused by non-modifiable hardware scheduling after tape-out. Conduct modeling and analysis on the computing power, bandwidth and power consumption of LPU chips. Carry out architecture performance benchmarking and optimization by comparing with traditional GPU and NPU architectures. Track the inference requirements of MoE large models and multimodal large models, iterate and upgrade the LPU hardware architecture to adapt to next-generation large model inference scenarios. Participate in chip front-end design and FPGA prototype verification, cooperate with the back-end team to complete chip tape-out, and follow up chip bring-up testing and performance optimization. Research dataflow chip architectures of overseas benchmark manufacturers including Groq, Etched and Cerebras, and deliver competitive analysis reports and architecture iteration solutions. Qualifications Master’s degree or above in Microelectronics, Integrated Circuit, Computer Architecture or related majors, with no less than 2 years of working experience in AI chip architecture design. Familiar with static dataflow architecture and systolic array architecture; understand the principles of Prefill and Decode dual-stage inference for large models. Prior experience in video memory and on-chip SRAM scheduling is preferred. Hands-on experience in AI chip / NPU / GPU architecture design, familiar with chip front-end design flow, with solid hardware-software co-design capabilities. Proficient in chip performance evaluation methodologies, capable of independently completing simulation and analysis of computing power, latency and power consumption. Preferred Skills R&D experience in dataflow chips or LPU chips. Experience in hardware adaptation for large model inference chips. Familiar with the fundamentals of AI compilers. In-depth research experience on Groq architecture.