Home Job Details
N
Information Technology 🏢 Full Time ⭐️ Verified

Senior LLM Engineer - Architecting the Future of Intelligence

Nexus Horizon AI
San Francisco
Estimated Salary
USD 180.000 – USD 250.000
Live Update
28 Juni 2026
Deadline
28 Jun 2027

Job Description

Be the Architect of the 2026 AI Landscape

We are not just building software; we are architecting the intelligence of tomorrow. Nexus Horizon AI is seeking a visionary Senior LLM Engineer to lead our generative AI initiatives. As we prepare for the next era of artificial general intelligence, we need a technical leader who thrives on complexity and innovation. You will define the architecture for next-generation Large Language Models, pushing the boundaries of what is possible in 2026 and beyond.

Why Join Us?

  • Work on cutting-edge projects that define the future of human-AI interaction.
  • Competitive compensation and equity packages for top-tier talent.
  • Flexible remote-first culture with a hub in the heart of SF.
  • Access to the latest GPU clusters and research infrastructure.

Key Responsibilities

  • Design, train, and deploy scalable Large Language Models (LLMs) for production environments.
  • Optimize model inference latency and reduce token costs through advanced quantization and pruning techniques.
  • Lead the implementation of Retrieval-Augmented Generation (RAG) pipelines to enhance model accuracy.
  • Collaborate with data scientists to curate high-quality synthetic training data sets.
  • Conduct cutting-edge research to improve prompt engineering strategies and context window management.
  • Mentor junior engineers and foster a culture of technical excellence and innovation.

Qualifications

  • Master’s or Ph.D. in Computer Science, Machine Learning, or a related field.
  • 5+ years of experience in deep learning, NLP, or Generative AI.
  • Expert proficiency in Python, PyTorch, and TensorFlow.
  • Strong experience with Hugging Face Transformers and model fine-tuning (LoRA, QLoRA).
  • Demonstrated success in deploying LLMs in high-volume production environments.
  • Experience with vector databases (Pinecone, Milvus, Weaviate) and cloud infrastructure (AWS/GCP).

Responsibilities

  • Design, train, and deploy scalable Large Language Models (LLMs) for production environments.
  • Optimize model inference latency and reduce token costs through advanced quantization and pruning techniques.
  • Lead the implementation of Retrieval-Augmented Generation (RAG) pipelines to enhance model accuracy.
  • Collaborate with data scientists to curate high-quality synthetic training data sets.
  • Conduct cutting-edge research to improve prompt engineering strategies and context window management.
  • Mentor junior engineers and foster a culture of technical excellence and innovation.

Qualifications

  • Master’s or Ph.D. in Computer Science, Machine Learning, or a related field.
  • 5+ years of experience in deep learning, NLP, or Generative AI.
  • Expert proficiency in Python, PyTorch, and TensorFlow.
  • Strong experience with Hugging Face Transformers and model fine-tuning (LoRA, QLoRA).
  • Demonstrated success in deploying LLMs in high-volume production environments.
  • Experience with vector databases (Pinecone, Milvus, Weaviate) and cloud infrastructure (AWS/GCP).

Required Skills

Python PyTorch TensorFlow NLP Hugging Face LLM Fine-tuning RAG Machine Learning Deep Learning AWS GCP

Ready to Take This Challenge?

Make sure your resume is ready. Submit your application now before the deadline.

Apply Now

Related Jobs

Similar job recommendations for you

View All