Job Description
Join the architects of tomorrow. Nebula Future Systems is pioneering the next evolution of artificial intelligence, and we are looking for a visionary Senior AI Infrastructure Architect to design the robust, scalable backbone of our global neural networks.
In this pivotal role, you won't just manage servers; you will engineer the physical and digital environments that will power the year 2026 and beyond. You will bridge the gap between cutting-edge machine learning research and high-performance computing infrastructure.
Why Nebula Future Systems?
We are a collective of futurists, engineers, and dreamers building the infrastructure for AGI. We offer a competitive compensation package, equity in a unicorn startup, and the opportunity to define the industry standard for AI compute.
Responsibilities
- Architect Scalable AI Workloads: Design and deploy high-availability, distributed AI training and inference pipelines capable of handling exascale data.
- Infrastructure Modernization: Lead the migration to next-gen cloud-native environments, optimizing for quantum-ready architectures and edge computing nodes.
- Performance Engineering: Continuously monitor, optimize, and tune system performance to reduce latency and maximize computational efficiency.
- Cost Optimization: Implement FinOps strategies to manage cloud resource expenditures while maintaining peak performance standards.
- Security & Compliance: Spearhead security initiatives to protect proprietary algorithms and data sovereignty across global regions.
- Collaborative Innovation: Work closely with research scientists and data engineers to translate theoretical models into production-ready infrastructure.
Qualifications
- Experience: 7+ years of experience in systems architecture, DevOps, or site reliability engineering, with a specific focus on AI/ML workloads.
- Technical Stack: Proficiency in Python, Go, or Rust, and deep expertise in Kubernetes, Docker, and AWS or GCP.
- AI Knowledge: Strong understanding of deep learning frameworks (TensorFlow, PyTorch) and how they interact with distributed systems.
- Problem Solving: Exceptional ability to troubleshoot complex, multi-layered system failures under pressure.
- Leadership: Proven track record of leading technical teams and mentoring junior engineers in architectural best practices.
- Communication: Ability to articulate complex technical concepts to non-technical stakeholders and executive leadership.