Job opening

Senior Manager, AI Platform Architecture

Calance

Bolingbrook, IL

Filed under IT Services and IT Consulting

Full job description

We have a full time opportunity with one of our major clients. They are looking for a Senior Manager, AI Platform Architecture for an important project.


Position: Senior Manager, AI Platform Architecture

Location: Bolingbrook , IL 60440, (Hybrid in Bollingbrook, IL – Tues, Wed, Thurs – every other month)

Duration: Full Time : Permanent


Job Description:

Agentic AI platform design and architecture

Multi-agent orchestration patterns

State and memory management approaches

LLM-as-a-Judge frameworks


Ideal Candidate:

Senior Manager, AI Platform Engineering or AI Platform Architect with 12–18 years of experience building enterprise AI/ML platforms. Strong background in Databricks, cloud architecture, MLOps, platform engineering, AI governance, and leading teams of engineers and architects. Experience supporting AI model development and deployment at scale while partnering with data science, product, security, and infrastructure team


What we're looking for

We need candidates who can go beyond strategy and team leadership and speak in detail about the architecture and implementation of enterprise AI platforms. Now do we need them to be able to go deep on everything below? No, that’s not realistic, but I hope this helps to paint a better picture of what to target in future conversations.


The strongest candidates should be able to discuss:

Agentic AI platform design and architecture

Multi-agent orchestration patterns

State and memory management approaches

LLM-as-a-Judge frameworks

MCP (Model Context Protocol) servers and agent integration frameworks

RAG architectures, context management, and knowledge services

Semantic layer strategy and tooling

Human-in-the-loop workflows

Prompt management and agent lifecycle/versioning

AI platform governance and operational controls

Security & Governance Depth

Candidates should be able to describe:

AI permissions and security strategies

Identity and access management approaches

Multi-agent security frameworks

Strategies for securing sensitive data in LLM environments

Model Armor, guardrails, and enterprise AI controls

Compliance, auditability, and responsible AI practices

AI Platform Operations & Observability


We're specifically looking for leaders who have personally driven or architected:

MLOps / AIOps frameworks

Logging and monitoring pipelines

Agent and model observability

Cost observability and optimization

Datadog and/or similar observability platforms

Continuous training and deployment pipelines

CI/CD processes for AI platforms

Cloud & Platform Architecture

The ideal candidate should be able to discuss trade-offs and design decisions across:

Vertex AI / GCP

Databricks

OpenAI ecosystem

Gemini ecosystem

Build vs. buy decisions

Platform selection criteria

Enterprise-scale AI infrastructure design


Example of the level of detail we're seeking

Rather than saying:

"I led an AI platform team that built agents."

We'd expect candidates to be able to explain:

"We standardized on Vertex AI with LangGraph for orchestration, implemented a multi-agent architecture with shared memory services, used RAG backed by Databricks vector search, integrated an MCP layer for tool connectivity, implemented evaluation using LLM-as-a-Judge frameworks, and monitored agent performance and cost through Datadog and custom observability dashboards."

Description:

Senior Manager, AI Platform Architecture


What we're looking for

We need candidates who can go beyond strategy and team leadership and speak in detail about the architecture and implementation of enterprise AI platforms. Now do we need them to be able to go deep on everything below? No, that’s not realistic, but I hope this helps to paint a better picture of what to target in future conversations.


The strongest candidates should be able to discuss:

Agentic AI platform design and architecture

Multi-agent orchestration patterns

State and memory management approaches

LLM-as-a-Judge frameworks

MCP (Model Context Protocol) servers and agent integration frameworks

RAG architectures, context management, and knowledge services

Semantic layer strategy and tooling

Human-in-the-loop workflows

Prompt management and agent lifecycle/versioning

AI platform governance and operational controls

Security & Governance Depth

Candidates should be able to describe:

AI permissions and security strategies

Identity and access management approaches

Multi-agent security frameworks

Strategies for securing sensitive data in LLM environments

Model Armor, guardrails, and enterprise AI controls

Compliance, auditability, and responsible AI practices

AI Platform Operations & Observability


We're specifically looking for leaders who have personally driven or architected:

MLOps / AIOps frameworks

Logging and monitoring pipelines

Agent and model observability

Cost observability and optimization

Datadog and/or similar observability platforms

Continuous training and deployment pipelines

CI/CD processes for AI platforms

Cloud & Platform Architecture


The ideal candidate should be able to discuss trade-offs and design decisions across:

Vertex AI and GCP

Databricks

OpenAI ecosystem

Gemini ecosystem

Build vs. buy decisions

Platform selection criteria

Enterprise-scale AI infrastructure design

Example of the level of detail we're seeking

Rather than saying:

"I led an AI platform team that built agents."

We'd expect candidates to be able to explain:

"We standardized on Vertex AI with LangGraph for orchestration, implemented a multi-agent architecture with shared memory services, used RAG backed by Databricks vector search, integrated an MCP layer for tool connectivity, implemented evaluation using LLM-as-a-Judge frameworks, and monitored agent performance and cost through Datadog and custom observability dashboards."


Going forward, we'd appreciate candidates who have demonstrable hands-on architecture experience in enterprise AI platforms, not primarily people leadership or program oversight. We are specifically looking for leaders who can speak in detail about agentic architectures, MLOps/AIOps, AI governance, platform security, observability, Vertex AI/Databricks ecosystems, and the technical design decisions behind those implementations. The ideal candidate should be comfortable operating at both the leadership level and the architectural implementation level.


Retail or eCommerce experience

Generative AI / LLM platform experience

Agentic AI frameworks and implementations

Snowflake, Kafka, Spark

Advanced observability and monitoring platforms

Multi-cloud experience

Responsible AI governance programs

AI cost optimization experience

Databricks certifications

Experience leading globally distributed teams


Nice to Have Skills:

Retail or eCommerce experience

Generative AI / LLM platform experience

Agentic AI frameworks and implementations

Snowflake, Kafka, Spark

Advanced observability and monitoring platforms

Multi-cloud experience

Responsible AI governance programs

AI cost optimization experience

Databricks certifications

Experience leading globally distributed teams


Technologies resource will be touching:

Databricks

GCP

AWS

Azure

AI/ML platforms

MLOps frameworks

CI/CD pipelines


Apply on original listing