Strong experience in designing and implementing high-performance, large-scale distributed systemsProven experience in implementing and deploying AI/ML platforms at scaleExpertise in building agent-based architectures, evaluation frameworks, and prompt/context engineeringKnowledge of MCP (Model Context Protocol) serversHands-on experience in LLM inference optimization, including batching and caching strategiesStrong experience with Kubernetes and cloud infrastructure (AWS/Azure/GCP)Proficiency in at least one programming language (Python, Java, Go, etc. )Expertise in designing agent data stacks & retrieval systems, including:Vector databasesHybrid searchData freshness strategiesMemory systemsGraph reasoningBM25 and advanced retrieval techniquesKey Responsibilities Architect and deliver scalable, high-performance distributed systemsDesign and deploy AI/ML and GenAI platforms at enterprise scaleBuild and manage agent-based architectures, including:Prompt and context engineeringMCP serversEvaluation frameworksOptimize LLM inference pipelines for latency, throughput, and efficiencyDesign and implement agent data & retrieval systems (vector DBs, hybrid search, memory, graph-based reasoning)Lead Kubernetes-based, cloud-native deploymentsProvide technical leadership, architecture governance, and hands-on mentoring to engineering teams
Create an account to see the full posting, access our search engine, and more.You're just 60 seconds away from your new Creativeloft account.