Tanggung Jawab Utama
Design and implement inference runtime components in Rust and C++ for ultra-low latency execution
Profile and optimize model execution across heterogeneous hardware targets (ARM Cortex-M55, Apple Silicon, NVIDIA Jetson, Tenstorrent)
Build adaptive quantization pipelines with automated accuracy benchmarking and layer sensitivity profiling
Collaborate with research scientists to translate frontier papers into deterministic production code
Write robust integration tests for sub-12ms latency-critical paths and memory-constrained execution
Build telemetry instrumentation to catch activation drift and thermal throttling in real time
Kualifikasi Wajib (Must-Have)
- 5+ years systems programming (Rust strongly preferred, C++ accepted)
- Deep understanding of CPU/GPU/NPU memory hierarchies and memory-constrained execution
- Demonstrated track record of profiling, optimizing, and deploying sub-15ms ML runtimes
- Familiarity with operator fusion, quantized graph representations, and memory-mapped models
Nilai Tambah (Nice-to-Have)
- + Experience with custom ML accelerator compilation (TVM, IREE, LLVM backends)
- + Open source contributions to high-performance inference engines
- + Background in robotics, aerospace, or automotive real-time systems
Kompensasi & Ekosistem
Membangun di Fantasma Synergy
Kepemilikan Substantif
Alokasi kepemilikan saham yang transparan untuk setiap engineer dan peneliti dengan jadwal vesting standar industri.
Akses Silikon Tanpa Batas
Kluster cloud TPU/GPU khusus serta hardware lab edge (NPU testbed, logic analyzers, oscilloscope, robotics rig).
Cakupan Global Lengkap
Asuransi kesehatan komprehensif tingkat dunia, tunjangan relokasi resmi ke Singapura atau Zurich, dan visa sponsorship.
Dukungan Konferensi Penuh
Biaya perjalanan dan registrasi tanpa batas untuk mempresentasikan karya di NeurIPS, ICML, ICLR, OSDI, dan SOSP.
Alur Seleksi Teknis (Zero Fluff)
Review Portofolio
Diskusi asinkron 30 menit mengenai kode atau paper yang pernah Anda bangun.
Deep Dive Arsitektur
Analisis desain sistem nyata 60 menit bersama tim engineer inti Fantasma Synergy.
Hands-on Pairing
Sesi coding terfokus pada profiling latensi, optimasi memori, atau kernel runtime.
Penawaran Resmi
Penyelarasan kompensasi, ekuitas, dan tanggal mulai kerja dalam 48 jam.
Formulir Aplikasi
Lamar untuk Research Engineer, Neural Fabric
Kami meninjau setiap aplikasi secara mendalam tanpa screening bot otomatis. Anda akan menerima kabar dalam 48–72 jam.