Job opening

Remote noted

Product Manager - GPUaaS and OE Telemetry

GMI Cloud

Mountain View, CA

Filed under IT System Data Services

Full job description

About us

GMI Cloud is a fast-growing, AI-native infrastructure company delivering high-performance GPU compute, inference services, and infrastructure for AI agents.


Following 8x ARR growth, GMI Cloud continues to scale rapidly across the U.S. and APAC. As a Reference Platform NVIDIA Cloud Partner (NCP) and a validated leading NCP across both markets, we power production AI for leading AI-native companies including Fireworks AI, Cartesia, Reflection, and OpenRouter.


From large-scale compute to optimized inference and agentic workloads, GMI Cloud gives AI teams the infrastructure they need to build, deploy, and scale on one unified cloud.


One cloud for compute, inference, and agents.


The Role:

We are seeking an experienced Product Manager to lead the strategy, roadmap, and execution of our GPU-as-a-Service (GPUaaS) platform and Observability (OE) platform. This role is responsible for defining products that enable customers to seamlessly consume GPU infrastructure while empowering engineering and operations teams with comprehensive observability across infrastructure, platform services, and AI workloads through telemetry, monitoring, analytics, and automation.


Preferred Location: Remote, USA.


Responsibilities

  1. Drive the product strategy and roadmap for the GPUaaS and OE Telemetry platform in collaboration with the Infrastructure Engineering team.
  2. Translate customer and engineering requirements into prioritized product roadmaps.
  3. Define product requirements, user stories, acceptance criteria, and success metrics.
  4. Lead product planning, roadmap reviews, and release planning.
  5. Measure product adoption, operational impact, and business outcomes using data-driven KPIs.
  6. Define capabilities requirements for GPU provisioning, lifecycle management, scheduling, orchestration, self-service portal, APIs, multi-tenancy, billing, metering, quotas, and access management.
  7. Define telemetry and observability capabilities requirements across GPU, compute, storage, networking, Kubernetes, Slurm, AI workloads, and supporting infrastructure.
  8. Drive capabilities adoption including (but not limited to):
  9. Metrics, logs, traces, and events collection
  10. OpenTelemetry adoption and instrumentation
  11. Real-time dashboards and visualization
  12. Intelligent alerting and incident detection
  13. Service health and dependency mapping
  14. Distributed tracing
  15. Root cause analysis
  16. Utilization analytics
  17. AI-driven anomaly detection and predictive insights
  18. SLO/SLI measurement and reliability reporting
  19. Partner with Infra Engineering teams, SRE, Platform Developers, Data Center Operations to improve observability, reliability, scalability and operation experience for GPUaaS platform.


Qualifications

  1. Bachelor's degree in Computer Science, Engineering, or a related technical field.
  2. 5+ years of Product Management experience in cloud infrastructure, AI infrastructure, or enterprise platforms.
  3. Strong understanding of GPU computing (NVIDIA H100/H200/B200/B300 or equivalent), Kubernetes, Slurm, Linux, Cloud infrastructure, APIs, Telemetry and observability platforms
  4. Experience with monitoring technologies such as Prometheus, Grafana, OpenTelemetry, Elasticsearch/OpenSearch, or similar.
  5. Experience translating customer requirements into technical product specifications.
  6. Strong analytical and communication skills.
  7. Experience building AI cloud or GPU cloud platforms.
  8. Knowledge of NVIDIA AI Enterprise and related software ecosystems NVSentinel, Fleet Intelligence, etc.
  9. Experience with multi-tenant IaaS platforms.
  10. Experience with Agile/Scrum product development.
  11. Meeting every qualification is not required—if you’re excited about this role, we’d love to hear from you. We believe diverse perspectives and experiences strengthen our team.


Apply on original listing