Systemadministrator (m/w/d)
Skills
About the role
About the Role We are building a highly secure, on-premise Large Language Model (LLM) infrastructure that operates entirely offline. We are looking for a AI Infrastructure Engineer who operates at the intersection of bare-metal hardware management and MLOps. In this role, you will be responsible for building, optimizing, and maintaining the physical and software architecture required to serve massive AI models securely and efficiently without relying on external cloud APIs.
Key Responsibilities
Bare-Metal GPU Management: Install, configure, and troubleshoot Linux servers, Nvidia GPU drivers, CUDA toolkits, and cuDNN versions to ensure maximum hardware utilization.
Local Inference Deployment: Deploy and manage high-performance local inference servers such as vLLM, Hugging Face TGI (Text Generation Inference), or TensorRT-LLM.
Model Optimization: Manage VRAM efficiently through model quantization (GGUF, AWQ, EXL2) and calculate hardware requirements for various model sizes and batch loads.
Containerization & Orchestration: Package model weights and inference engines into Docker containers and orchestrate deployments using Kubernetes (K8s).
API & Integration: Set up and maintain local reverse proxies (e.g., LiteLLM) to ensure the offline models provide a standard, OpenAI-compatible API for internal developers.
Security & Monitoring: Architect and maintain air-gapped network environments with zero outbound internet access. Implement robust monitoring using Prometheus and Grafana to track GPU temperatures, power limits, and VRAM usage.
Must-Have Qualifications
Extensive experience as a Linux System Administrator managing bare-metal servers.
Deep, hands-on understanding of Nvidia GPU architecture, multi-GPU orchestration (NVLink, NCCL), and resolving driver/CUDA conflicts.
Proven experience deploying open-source LLMs (Llama 3, Mistral, Mixtral) locally using frameworks like vLLM or Ollama.
Strong proficiency in containerization (Docker) and orchestration (Kubernetes).
Solid scripting skills in Python and Bash for automation and MLOps pipelines.
Strong understanding of network security, specifically designing and maintaining air-gapped systems.
Nice-to-Have
Experience with custom model fine-tuning infrastructure.
Familiarity with building retrieval-augmented generation (RAG) pipelines on local hardware.
Location - Onsite Start Date - Immediately Type - Mini job (10 hours a week)
Job Type: Part-time
Pay: Up to 13,90€ per hour
Work Location: In person
Questions about this role
Want AI Applyd to auto-apply to roles like this?
We tailor your resume per posting, fill the forms, and track replies for you.