June 6, 2026

Senior AI Site Reliability Architect

Oracle Jackson, Mississippi

Job Description

Join us as a Senior AI Site Reliability Architect and be a key player in the development and management of a state-of-the-art, AI-first Electronic Health Record platform. In this dynamic role, you will design, build, and maintain highly reliable, scalable infrastructure and data pipelines critical for global analytics.

You will drive the future of cloud operations by advancing automation, observability, and AI-assisted reliability practices. This includes experimenting with Generative AI and intelligent automation to enhance incident response, system resilience, and operational efficiency.

Collaborate with a talented team to deliver robust solutions that efficiently handle vast datasets while continuously enhancing system reliability and operational excellence.

Note: U.S. citizenship is required for this role, as the selected candidate will need to obtain and maintain a U.S. government security clearance.

Required Skills

Infrastructure & Reliability

Experience in building and operating high-availability, fault-tolerant systems. Strong understanding of distributed systems and resiliency patterns. Proven track record in incident response and production troubleshooting. AI-Native Engineering (NEW)

Hands-on experience with Generative AI or Agentic AI (e.g., LangChain, AutoGPT) in: Infrastructure lifecycle management. Observability and anomaly detection. Automated incident response and remediation. Designing AI-driven workflows for operational efficiency. Building autonomous agents for DevOps/SRE use cases. Cloud & Multi-Cloud Ecosystems

Extensive experience in multi-cloud environments (OCI, AWS/Azure). Deep understanding of cloud infrastructure design and resource optimization. Management experience with hybrid or cross-cloud architectures. DevOps/SRE Practices

Advanced skills in CI/CD pipelines (Jenkins, Kubernetes). Proficiency with Infrastructure as Code (Terraform). Experience with observability tools (Prometheus, Grafana). Strong emphasis on automation-first operations. Data Technologies

Proficient in data warehousing platforms (e.g., Vertica, Snowflake), ETL frameworks, and large-scale data processing.

BI & Reporting

Experience integrating BI tools (Tableau, Power BI, Oracle Analytics).

Programming & Tools

Strong proficiency in Python, Java, or Go. Experience with Docker, Kubernetes, and shell scripting. Problem-Solving

Excellent troubleshooting abilities with a focus on root-cause analysis. Experience resolving complex issues in distributed systems. Responsibilities

Collaborate with the Site Reliability Engineering (SRE) team to take shared ownership of service and platform components, understanding end-to-end system architecture and dependencies. Design, build, and operate reliable, scalable, and secure infrastructure for large-scale analytics workloads. Enhance system reliability through automation, monitoring, and continuous performance optimization. Embrace AI-assisted approaches for operations, such as: Improving observability and alerting. Supporting automated incident detection and remediation. Exploring intelligent automation for infrastructure lifecycle management. Partner with development teams to enhance architecture, scalability, and operability. Participate in on-call rotations to tackle complex production issues. Perform root cause analysis and implement solutions to prevent recurrence. Leverage knowledge of distributed systems to optimize performance. Drive continuous improvement in DevOps/SRE practices, including CI/CD, Infrastructure as Code, and automation. Develop & Maintain

Implement and optimize infrastructure for Oracle HDI Analytics Platform. Ensure system uptime and scalability. AI-Driven Automation (NEW)

Design and implement GenAI-powered or agent-based solutions for: Observability and anomaly detection. Incident triage and remediation. Infrastructure provisioning and management. Building tools for self-service and autonomous operations. Data Pipeline Execution

Build and optimize scalable data pipelines using Vertica and ETL frameworks. Operational Excellence

Utilize DevOps/SRE practices to automate operations. Enhance observability with Prometheus/Grafana and AI insights. Cloud Integration

Support multi-cloud initiatives across OCI, AWS, and Azure. Optimize performance, cost, and compliance across environments. Incident Response

Participate in on-call rotations. Implement automated remediation solutions. Collaboration

Work closely with engineers to execute technical initiatives. Contribute to code reviews and infrastructure enhancements. What You Bring

4+ years of experience in software engineering, cloud infrastructure, or SRE. Proven record of ensuring production system reliability. Core Expertise

Cloud infrastructure design and automation. Distributed systems and performance optimization. Data warehousing and ETL frameworks. AI-Native Experience

Experience applying GenAI/LLMs to infrastructure or operations. Proven capability in AI-powered automation for DevOps/SRE. Familiarity with tools such as LangChain, AutoGPT, or custom AI agents. Technical Skills

Terraform, Docker, Kubernetes. Observability stacks (Prometheus, Grafana). Proficiency in Python, Java, or Go. Additional Strengths

Problem-solving mindset focused on automation. Experience improving reliability through intelligent automation. Preferred Qualifications

Experience in regulated environments (HIPAA, compliance frameworks). Experience in environments requiring security clearance. Experience with self-healing or autonomous infrastructure systems. This role may require compliance with certain requirements, including potential immunization mandates.

Hiring Range: $79,200 to $178,100 per annum, with eligibility for bonus and equity.

Oracle offers a comprehensive benefits package that includes medical, dental, and vision insurance, among other benefits.

Discover your potential at a company leading the AI and cloud solutions industry, impacting lives globally.

Oracle is an Equal Employment Opportunity Employer, committed to inclusion and accessibility.

Create a free account to keep reading — and to apply.

Create My Free Account

TheCreativeLoft is a better way to find jobs. Find out more:

You're just 60 seconds away from your new Creativeloft account.

Looking For Similar Jobs?