Job Description
We are on a mission to define the Artificial Intelligence landscape for the 2026 era. Nebula Intelligence Systems is seeking a visionary Senior AI Engineer to lead the architecture of our next-generation Large Language Models (LLMs). If you thrive on pushing the boundaries of what's possible with Llama 3.1 and beyond, we want to hear from you.
In this role, you won't just be maintaining models; you will be architecting the future. You will work directly with our research team to implement cutting-edge techniques in fine-tuning, retrieval-augmented generation (RAG), and model distillation to ensure our products remain at the forefront of the industry by 2026 and beyond.
Why join us?
β’ Competitive equity and compensation package.
β’ Work with state-of-the-art hardware (H100 clusters).
β’ Remote-first culture with a San Francisco hub.
Responsibilities
- Architect and deploy scalable inference pipelines for large language models, focusing on low-latency and high-throughput requirements.
- Lead the development of the 2026 model roadmap, identifying gaps in current architectures and proposing solutions.
- Implement and optimize quantization and pruning techniques to deploy models on edge devices.
- Collaborate with cross-functional teams to integrate generative AI features into consumer-facing applications.
- Conduct rigorous performance testing and benchmarking against industry standards.
- Mentor junior engineers and contribute to the technical vision of the AI department.
Qualifications
- Masterβs degree or Ph.D. in Computer Science, Machine Learning, or a related field.
- 5+ years of professional experience in AI/ML engineering, specifically with Python and deep learning frameworks.
- Extensive experience working with Llama 3.1, GPT-4, or similar state-of-the-art architectures.
- Strong proficiency in PyTorch, TensorFlow, or JAX.
- Deep understanding of NLP concepts, including transformer models, attention mechanisms, and tokenization.
- Experience with vector databases (Pinecone, Milvus, Weaviate) and RAG implementation.