Trainium (NKI) Kernel Expert
About this role
Trainium (NKI) Kernel Expert
Remote | Independent Contractor | United States | $145,600–$187,200 annualized ($70–$90/hour)
About the Role
Apply your specialized AWS Trainium and Neuron Kernel Interface (NKI) expertise to help improve the quality and reliability of advanced AI systems.
As a Trainium (NKI) Kernel Expert, you’ll evaluate NKI development tasks used to train and assess frontier AI models. You’ll assess whether kernels are technically correct, appropriately designed for Trainium hardware, and optimized for the platform.
Your expertise will be particularly valuable when evaluating CUDA-to-NKI migrations, Trainium-specific performance optimization, and cross-platform numerical correctness.
What You’ll Do
Evaluate NKI kernel development tasks for quality, correctness, completeness, and hardware appropriateness.
Assess the fidelity and technical quality of CUDA-to-NKI migrations.
Evaluate Trainium-specific kernel optimization strategies and performance considerations.
Review tile-based computation approaches and NKI memory-management patterns.
Assess the appropriate use of SBUF, PSUM, and HBM within kernel implementations.
Evaluate partition-dimension constraints and DMA orchestration.
Assess Trainium kernel performance, including:
NeuronCore pipeline utilization
Tensor-engine throughput
Memory-bandwidth bottlenecks
Evaluate cross-platform numerical correctness between GPU and Trainium implementations.
Assess differences arising from accumulation order, rounding behavior, and mixed-precision semantics.
Provide clear, structured, rubric-based written feedback on kernel tasks and solutions.
What You Bring
2+ years of hands-on experience developing or optimizing kernels with Neuron Kernel Interface (NKI) for AWS Trainium or Inferentia2 hardware.
Strong understanding of NKI development patterns, including:
Tile-based computation
SBUF, PSUM, and HBM memory hierarchy
Partition-dimension constraints
DMA orchestration
Demonstrated experience assessing or performing CUDA-to-NKI kernel migrations.
Familiarity with Trainium-specific performance profiling and optimization.
Strong understanding of cross-platform numerical correctness, including:
Accumulation order
Rounding behavior
Mixed-precision semantics
Ability to analyze technical implementations with precision and communicate findings clearly in writing.
Preferred Qualifications
Experience in the following areas is highly valuable:
Hands-on experience with the AWS Neuron SDK.
Familiarity with Neuron Compiler internals.
Contributions to NKI kernel libraries or related projects.
Previous CUDA or Triton kernel development experience.
Strong understanding of Trainium hardware architecture and capabilities, including:
NeuronCore-v2 architecture
On-chip SRAM topology
FP32, BF16, FP8, and INT8 data types
Experience benchmarking machine-learning training workloads on Trn1 or Trn2 instances.
Compensation & Engagement
Rate: $70–$90/hour
Annualized Equivalent: $145,600–$187,200
Location: United States
Work Arrangement: Fully remote
Engagement Type: Independent contractor
Schedule: Flexible
Payment: Weekly via Stripe or Wise
Annualized compensation is based on 2,080 hours per year for comparison purposes only. Actual earnings depend on the number of hours and projects completed.
Why This Opportunity?
Apply your Trainium and NKI expertise to advanced AI development.
Work on technically challenging kernel evaluation and optimization problems.
Help improve AI systems' understanding of specialized accelerator architectures.
Evaluate real-world kernel migration, performance, and numerical-correctness challenges.
Contribute expertise in an increasingly important AI hardware ecosystem.
Work remotely with a flexible schedule.
Earn competitive compensation for highly specialized technical expertise.
Contract & Payment Terms
You will be engaged as an independent contractor.
Work is fully remote and can be completed on your own schedule.
Projects may be extended, shortened, or concluded early depending on project needs and performance.
Your work will not require access to confidential or proprietary information belonging to any employer, client, or institution.
Payments are made weekly via Stripe or Wise based on services rendered.
H-1B and STEM OPT candidates cannot be supported at this time.
Equal Opportunity
All qualified applicants will be considered without regard to legally protected characteristics. Reasonable accommodations are available upon request.
- Fully remote contract role open to candidates in United States.
- Compensation: $145,600 - $187,200/year, paid in USD.
- Vetted and managed by Recruitment Room — no placement fees for candidates.
- Flexible hours; part-time and full-time engagements available.
More CUDA roles
- Senior Software EngineerMultiple countries · $100 - $150/hour
- Senior Software EngineerMultiple countries · $30 - $100/hour
- Product & Engineering ExpertMultiple countries · $288,000 - $672,000/year
- Backend Security EngineerMultiple countries · $57,600 - $192,000/year
- VP of Robotics ResearchUnited States · $250,000 - $350,000/year
- Computer User Support SpecialistMultiple countries · $30 - $55/hour