Nvidia logo
英伟达
NCX Engineer, AI Accelerator

NCX Engineer, AI Accelerator

发布于 大约 15 小时前

普通员工/个人贡献者

上海市 / 北京市
高级经验
全职员工
混合式弹性办公
本科
软件工程
Ai加速器
Go
Gpu
Mlops
Nvidia
Pytorch
Tensorflow
分布式训练
容器化

AI 估算 · 40k–70k

英伟达为AI领域顶尖公司,职位要求高经验和技术深度,薪资对标一线大厂高级工程师,预估月薪40-70K。

职位详情

关于这个职位

作为英伟达AI加速器团队的NCX工程师,你将与全球顶尖AI公司合作,为他们提供深度技术支持

你将负责构建和部署定制化的AI解决方案,优化分布式训练和推理性能,并确保客户能够充分发挥英伟达AI平台的潜力
这是一个深入前沿AI技术、与行业巨头并肩作战的绝佳机会

最低要求

BS, MS, or Ph.D. in Computer Science, Computer/Electrical Engineering, or a related technical field, or equivalent experience.

+ years of experience in customer facing technical roles such as Solutions Engineering, DevOps, Site Reliability, or ML Infrastructure Engineering, ideally supporting large‑scale cloud or service provider environments.
Strong expertise in Linux systems, distributed computing, Kubernetes, containers, and GPU scheduling on multi-tenant or service-provider platforms.
Demonstrated AI/ML experience supporting large‑scale training and inference workloads (e.g., LLMs, generative models, recommendation systems) in production or critically important environments.
Solid programming skills in Python/Go, with hands‑on experience using frameworks such as PyTorch or TensorFlow for training and serving.
Demonstrated capability to collaborate with customer and partner engineering teams in fast-paced environments, guide intricate technical investigations, and bring issues to root cause and resolution.
Excellent communication and technical presentation skills, with the ability to clearly articulate architectures, trade‑offs, and recommendations to both engineering and leadership audiences.

工作职责

Build and deploy custom AI solutions on NCP and Neo Cloud platforms, including distributed training, inference optimization, and MLOps pipelines constructed on NVIDIA reference architectures.

Act as the main technical contact for strategic NCPs, offer remote and on-site support, troubleshoot complex production problems, and guide partner engineering teams on NVIDIA platform guidelines.
Deploy and manage AI workloads across DGX Cloud, NCP data centers, and major CSP environments using Kubernetes, containers, and GPU scheduling systems aligned to NCP builds.
Profile and tune large-scale training and inference workloads on NCP platforms. Implement observability and SLO/SLA monitoring. Lead detailed efforts to reduce latency, cost, and operational risk.
Implement and expand NVIDIA reference architectures on partner platforms, develop integrations with partner control planes and customer environments, and ensure smooth API, data pipeline, and enterprise software connectivity.
Build detailed implementation guides, runbooks, and post‑mortem documentation that codify standard methodologies for running NVIDIA AI workloads at scale on NCP platforms.

优先资格

Experience with the NVIDIA ecosystem, including DGX systems, CUDA, NeMo, Triton, NIM, and NVIDIA networking technologies such as InfiniBand and RoCE.

Direct experience collaborating with NVIDIA Cloud Partners, hyperscale CSPs, or managed AI cloud platforms, including implementation of NVIDIA reference architectures for AI infrastructure.
Deep familiarity with MLOps and cloud‑native practices: containerization, CI/CD pipelines, observability stacks (Prometheus, Grafana, OpenTelemetry), and GitOps workflows.
Background in infrastructure as code (Terraform, Ansible, or similar) for repeatable deployment and configuration of GPU‑accelerated clusters and NCP building blocks.

AI 洞察

优缺点分析

优点

  • 身处AI行业最前沿,接触最先进的AI加速技术和架构
  • 与全球顶级AI公司合作,积累宝贵的人脉和行业经验
  • 英伟达平台的技术积累和资源支持,可不断提升自身技术深度
  • 需要频繁出差和客户现场支持,对工作和生活平衡有一定影响
  • 技术更新快速,需要持续学习以跟上AI基础设施的发展

缺点 / 挑战

  • 客户问题可能非常复杂且紧急,需在压力下快速排查和解决
  • 这个职位适合那些对AI基础设施充满热情、喜欢解决挑战性问题、并愿意在快节奏环境中与客户紧密合作的技术专家

角色解读

  • 成为AI基础设施领域的专家,主导客户的核心AI平台落地
  • 向解决方案架构师或技术总监方向发展,负责更大规模的客户项目
  • 深入NVIDIA生态,参与新一代AI加速技术的研发和推广
  • 与战略客户合作,为他们提供AI加速器相关的技术支持,解决生产环境中的复杂问题
  • 在NVIDIA云平台和合作伙伴平台上构建、部署和优化分布式训练和推理工作负载
  • 使用Kubernetes和GPU调度系统管理大规模AI工作负载,并实施监控和性能调优
  • 精通Linux、分布式计算、Kubernetes、容器化和GPU调度
  • 熟练掌握Python或Go,并有PyTorch或TensorFlow的实际使用经验
  • 具备8年以上面向客户的技术角色经验,能够处理大规模AI系统
  • 优秀的沟通和演示能力,能够向工程和领导层清晰传达技术方案

申请策略

  • 在简历中量化你过去项目的规模(如GPU节点数、训练吞吐量、延迟优化百分比)
  • 准备一些具体的故事,说明你如何通过技术方案帮助客户实现性能突破
  • 突出你在大规模分布式训练和推理方面的实战经验,特别是使用Kubernetes和GPU调度
  • 强调你与外部客户或合作伙伴的合作经历,以及如何解决复杂生产问题
  • 展示你在Python/Go和AI框架(PyTorch/TensorFlow)方面的编程能力
  • 补充英伟达生态相关的技能,如CUDA、NeMo、Triton、InfiniBand等
  • 学习MLOps和云原生工具链,包括Prometheus、Grafana、CI/CD和GitOps
  • 了解基础设施即代码(Terraform、Ansible)以提升自动化部署能力

面试指南

  • 使用STAR方法:描述情境、任务、行动和结果,突出你的技术贡献和客户影响
  • 对于技术问题,展示系统性思维:从问题定义、根因分析到方案实施和验证
  • 涉及团队协作时,强调你的沟通策略和跨部门协调能力
  • 请描述你设计和部署一个大规模分布式训练系统的经验
  • 当客户遇到GPU利用率低下的问题,你会如何排查和优化?
  • 谈谈你在Kubernetes上管理AI工作负载时遇到的一个挑战
  • 你如何与客户工程团队有效沟通复杂的技术方案?
  • 解释一下MLOps在AI基础设施中的作用,以及你如何实现CI/CD管道

职位点评

74
综合评分

英伟达AI加速器工程师,前沿技术栈,高成长性,薪资优厚但未披露,混合办公,有出差需求。

从薪资福利、成长空间、工作节奏和岗位方向综合评估,方便横向比较。

更适合这类人
最适合追求技术前沿和职业成长,愿意接受一定灵活性和工作强度的求职者。
表现最好
成长发展
相对薄弱
工作生活
薪资福利70
成长发展90
工作生活60
使命价值75

薪资福利

70中等

职位未明确薪资,但英伟达作为行业领先公司,整体薪酬具有较强竞争力,福利体系完善,因此补偿性动机满足程度中等偏上。

薪资信号未披露(AI估算:40K-70K/月)

成长发展

90较高

该职位涉及最前沿的AI加速技术,与顶级客户合作,技术挑战大,成长空间广阔,能够极大提升个人在AI基础设施领域的专业能力。

技术前沿前沿/新兴技术
技术栈AI加速器、分布式训练、Kubernetes、GPU、LLM、MLOps
业务类型ambiguous

工作生活

60中等

混合办公模式提供一定灵活性,但需要出差和客户现场支持,工作强度可能较大,生活化动机满足程度一般。

工作模式混合式弹性办公
办公地点市区核心地段
加班情况未提及(无法判断)

使命价值

75中等

AI行业属于高速增长赛道,职位对推动AI技术落地有直接贡献,社会意义中性偏正面,意义感动机得到较好满足。

行业发展高速增长赛道
社会影响中性/一般
创新程度积极采用新技术
Watch Jobs