Job

Staff Machine Learning Engineer

ServiceNow United States

Deadline: Not specified

Short Summary

Staff Machine Learning Engineer with ServiceNow. Team Overview We build the AI layer of our CPQ platform — a set of Python services that let users configure, quote, and manage transactions through natural language instead of forms. This isn't a thin LLM wrapper. We're running multiple production agent architectures concurrently (ReAct-style tool-calling agents, hand-rolled LangGraph state machines, and the Harness — our…

Key Details

  • Position / opportunity: Staff Machine Learning Engineer
  • Organization: ServiceNow
  • Country / coverage: United States
  • Location: Santa Clara, California, United States
  • Work arrangement: On-site
  • Opportunity type: Jobs
  • Sector: Education
  • Compensation: Not specified
  • Duration: Not specified
  • Start date: 2026-08-20
  • Application deadline: Not specified

Description

Team Overview We build the AI layer of our CPQ platform — a set of Python services that let users configure, quote, and manage transactions through natural language instead of forms. This isn't a thin LLM wrapper. We're running multiple production agent architectures concurrently (ReAct-style tool-calling agents, hand-rolled LangGraph state machines, and the Harness — our from-scratch, industry-leading agent execution runtime). Our systems are backed by a first-party MCP surface into admin/product/rules/transaction systems and interoperate with other AI agents over the A2A protocol. Below that sits a conventional Java/Spring Boot microservices fleet and a React/TypeScript + Lit frontend that the agents ultimately drive. Role Overview We're looking for someone who already operates at a Senior-Staff bar in the agentic/LLM domain but is building out breadth across the rest of the stack. You'll be one of the most senior technical voices on how agentic systems get designed here — state management, tool boundaries, streaming protocols, prompt/context architecture, and multi-agent coordination — while staying credible end-to-end: able to read a Spring Boot service, unblock a frontend integration, or reason about a classical ML model pipeline when the problem calls for it. What you get in this role: Multi-agent orchestration — LangGraph/LangChain agents over frontier LLMs for transaction editing, conversational configuration, and multi-product quote planning with plan/approve/refine loops and parallel task execution The Harness — we're crystallizing our own agent execution runtime into an industry-leading, state-of-the-art harness. Full-duplex sessions where a user can interrupt, redirect, or answer a clarifying question mid-execution while other work keeps streaming, built on a from-scratch async runtime rather than a bolted-on wrapper around someone else's agent loop. This is as much a performance and UX problem as a backend one — low-latency streaming, backpressure, live progress, partial results, graceful cancellation — and it's the part of the stack we're most invested in owning outright. You'd be a primary owner of where this goes next. MCP as a secondary interface — we maintain a first-party MCP server and clients into our admin/product/rules/transaction systems, but as the Harness matures it becomes the primary way our own agents interact with the platform, with MCP kept as the secondary, standards-based surface for external interop. You'd help decide what stays MCP-first and what moves onto the Harness. A2A protocol — agent-to-agent task delegation and streaming, surfaced through an external gateway so other systems (including core ServiceNow) can drive our agents directly Forward Deployed Engineering — expect real time embedded with customer- and product-facing teams against live deployments. Adapting the Harness and our agents to actual customer workflows under real constraints, not just building platform capability in the abstract RAG / context engineering — tenant-uploaded document ingestion, categorization, and aggregation into agent context. Prefix-cacheable prompt design for cost/latency Classical ML, when the problem isn't a good fit for an LLM — we have a separate PyTorch/scikit-learn training and serving pipeline (field-value prediction) that a whole-stack ML engineer should be able to read, extend, or evaluate against LLM-based alternatives Full-stack fluency — enough comfort in Spring Boot/Java services and the React/TypeScript + Lit frontend to unblock an integration end-to-end without waiting on a handoff To be successful in this role you have: 8+ years building production software, with several years specifically shipping LLM-powered / agentic systems (not just API wrapper calls — real tool-use loops, state management, multi-turn orchestration) Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving. This may include using AI-powered tools, automating workflows, analyzing AI-driven insights, or exploring AI's potential impact on the function or industry. Deep, hands-on expertise with LangGraph and/or LangChain (or the judgment to…

Responsibilities

  • Team Overview We build the AI layer of our CPQ platform — a set of Python services that let users configure, quote, and manage transactions through natural language instead of forms.
  • We're running multiple production agent architectures concurrently (ReAct-style tool-calling agents, hand-rolled LangGraph state machines, and the Harness — our from-scratch, industry-leading agent execution runtime).
  • Our systems are backed by a first-party MCP surface into admin/product/rules/transaction systems and interoperate with other AI agents over the A2A protocol.
  • Below that sits a conventional Java/Spring Boot microservices fleet and a React/TypeScript + Lit frontend that the agents ultimately drive.
  • Role Overview We're looking for someone who already operates at a Senior-Staff bar in the agentic/LLM domain but is building out breadth across the rest of the stack.
  • You'll be one of the most senior technical voices on how agentic systems get designed here — state management, tool boundaries, streaming protocols, prompt/context architecture, and multi-agent coordination — while staying credible end-to-end: able to read…

Requirements / Eligibility

  • with MCP kept as the secondary, standards-based surface for external interop.
  • You'd help decide what stays MCP-first and what moves onto the Harness.
  • A2A protocol — agent-to-agent task delegation and streaming, surfaced through an external gateway so other systems (including core ServiceNow) can drive our agents directly Forward Deployed Engineering — expect real time embedded with customer- and…
  • Adapting the Harness and our agents to actual customer workflows under real constraints, not just building platform capability in the abstract RAG / context engineering — tenant-uploaded document ingestion, categorization, and aggregation into agent context.

How To Apply

Use the Apply now button and follow the instructions on the application page.