Build the On-Device Future
We are building a telemetry-free local software ecosystem. Join our engineering lab to optimize local compilers, neural inference, and haptic designs.
Our Lab Principles
How we build next-generation applications at Headon AI Labs.
100% Privacy Base
We believe user inputs are sacred. We do not write code that logs, tracks, or uploads telemetry parameters. Every line must operate offline.
Latency is King
Every keystroke haptic or autocorrect predictive token must compute under 5ms. We optimize down to vector registers and compiler instruction levels.
Pure Interaction
Software is a tactile dialogue. We design responsive, physical haptics and glassmorphic micro-animations that make digital inputs feel alive.
Active Openings
Find your role in the offline computing revolution.
On-Device LLM Optimization Engineer
Role Overview
We are seeking a senior compiler/model optimization engineer to quantize and compile large language models to run directly on consumer mobile NPUs and CPUs. You will design custom inference loops targeting ARM NEON and Apple Neural Engine instructions.
Requirements
- Deep familiarity with deep learning quantization techniques (INT4, INT8, AWQ, GPTQ).
- Expertise in C++, CUDA, PyTorch, and WebGPU/WASM assembly.
- Experience writing custom execution kernels for mobile system enclaves.
Lead Interaction Designer (Haptics & Gestures)
Role Overview
You will lead the user experience design of our physical input layouts. This includes tuning linear resonant actuators (LRA) haptic wave parameters, designing micro-gestures, and writing beautiful glassmorphic interface guidelines.
Requirements
- Portfolio demonstrating exceptional tactile, motion, and interaction design systems.
- Deep knowledge of CSS animations, keyframes, SVG animation vectors, and WebGL.
- Ability to prototype haptic pulse sequences using native core-haptics APIs.
Embedded Systems Compiler Engineer
Role Overview
Help us write LLVM-level compiler extensions that translate high-level neural operations into raw, highly parallel machine instructions. Your goal is to maximize cache retention and avoid main memory access overhead during text predictions.
Requirements
- Expert level assembly compiler optimization experience.
- Familiarity with LLVM compiler infrastructure and target codegen configurations.
- Passionate about stripping telemetry layers and reducing executable binary weights.