Job Description
Be the Architect of the 2026 AI Landscape
We are not just building software; we are architecting the intelligence of tomorrow. Nexus Horizon AI is seeking a visionary Senior LLM Engineer to lead our generative AI initiatives. As we prepare for the next era of artificial general intelligence, we need a technical leader who thrives on complexity and innovation. You will define the architecture for next-generation Large Language Models, pushing the boundaries of what is possible in 2026 and beyond.
Why Join Us?
- Work on cutting-edge projects that define the future of human-AI interaction.
- Competitive compensation and equity packages for top-tier talent.
- Flexible remote-first culture with a hub in the heart of SF.
- Access to the latest GPU clusters and research infrastructure.
Key Responsibilities
- Design, train, and deploy scalable Large Language Models (LLMs) for production environments.
- Optimize model inference latency and reduce token costs through advanced quantization and pruning techniques.
- Lead the implementation of Retrieval-Augmented Generation (RAG) pipelines to enhance model accuracy.
- Collaborate with data scientists to curate high-quality synthetic training data sets.
- Conduct cutting-edge research to improve prompt engineering strategies and context window management.
- Mentor junior engineers and foster a culture of technical excellence and innovation.
Qualifications
- Master’s or Ph.D. in Computer Science, Machine Learning, or a related field.
- 5+ years of experience in deep learning, NLP, or Generative AI.
- Expert proficiency in Python, PyTorch, and TensorFlow.
- Strong experience with Hugging Face Transformers and model fine-tuning (LoRA, QLoRA).
- Demonstrated success in deploying LLMs in high-volume production environments.
- Experience with vector databases (Pinecone, Milvus, Weaviate) and cloud infrastructure (AWS/GCP).
Responsibilities
- Design, train, and deploy scalable Large Language Models (LLMs) for production environments.
- Optimize model inference latency and reduce token costs through advanced quantization and pruning techniques.
- Lead the implementation of Retrieval-Augmented Generation (RAG) pipelines to enhance model accuracy.
- Collaborate with data scientists to curate high-quality synthetic training data sets.
- Conduct cutting-edge research to improve prompt engineering strategies and context window management.
- Mentor junior engineers and foster a culture of technical excellence and innovation.
Qualifications
- Master’s or Ph.D. in Computer Science, Machine Learning, or a related field.
- 5+ years of experience in deep learning, NLP, or Generative AI.
- Expert proficiency in Python, PyTorch, and TensorFlow.
- Strong experience with Hugging Face Transformers and model fine-tuning (LoRA, QLoRA).
- Demonstrated success in deploying LLMs in high-volume production environments.
- Experience with vector databases (Pinecone, Milvus, Weaviate) and cloud infrastructure (AWS/GCP).