Job Overview
About Arbisoft:
Arbisoft is one of Pakistan’s most respected software engineering institutions, building world-class platforms for partners like edX, MIT, and leading Silicon Valley startups. We value clean code, engineering rigor, and deep problem solving.
Role Details:
Join our AI and Machine Learning systems division to engineer the infrastructure powering autonomous agents, large language model inference, and real-time retrieval-augmented pipelines.
Key Responsibilities:
• Design and maintain scalable infrastructure for distributed LLM inference, fine-tuning, and model evaluation.
• Build low-latency RAG systems combining vector databases, semantic search, and streaming token delivery.
• Optimize model serving latency and GPU resource utilization using vLLM, TensorRT-LLM, and Triton.
• Collaborate with academic researchers and product engineering teams on novel AI-assisted workflows.
• Implement robust guardrails, prompt evaluation frameworks, and telemetry for production agent interactions.
Requirements:
• 4+ years of software engineering experience with at least 2 years in Python and ML infrastructure.
• Solid understanding of transformer architectures, embeddings, vector indexing (Qdrant, Milvus, FAISS).
• Hands-on proficiency with Docker, Linux systems internals, and GPU-accelerated cloud instances (NVIDIA CUDA).
• Passion for building clean, well-tested systems with clear documentation.
