Member of the Technical Staff - Agentic Engineer
IT
San Francisco, CA, USA
Why Join Stand: At Stand, you’ll help build a new class of global property protection. We use advanced physics and AI to model catastrophic risk at the asset level, then automate underwriting and mitigation before loss occurs. Insurance is simply the current delivery mechanism. The real product is a scalable risk engine, our Stand World Model.
We stay when traditional insurers exit. We model what others approximate. And we build systems that change outcomes, not just prices.
Our leadership team includes former successful founders and CEOs from Metromile, PolicyGenius, WePay, and HotelTonight, bringing deep experience in building and scaling high-growth companies.
Background: The property insurance industry is built to price loss after it happens. It relies on coarse proxies, backward-looking data, and manual processes, then accepts damage as unavoidable.
Stand takes a different approach. We simulate how real-world catastrophes affect individual properties, translate that into actionable decisions, and automate the business around it. The result is a platform that can underwrite what others can’t and operate with far less friction.
Role Summary
We are building the self-driving insurer: a national carrier whose standard operating procedure is explicit, instrumented, and handed to agents one proven step at a time — so the book grows without the org.
That bet has a single point of failure, and this role owns it. The runtime, contracts, and eval framework that every agentic capability at Stand is built on. Not one agent. The framework that enabled the team to build and manage hundreds of them
Concretely, the harness is four things:
An eval framework and diagnostic center. Turning an existing data stream into an eval, or wiring up the right metric when one doesn't exist yet, should take an afternoon not a week. The eval framework should make building new ones delightful. The diagnostic center gives anyone, not just the builder, a high-level read on whether a skill is actually working.
A tooling center. Adding a new MCP capability, and making it permission-aware, should be easy to build and reuse. Deciding when someone's agent should be allowed to do X or Y isn't your call to make, it's your job to make it easy for an engineer to build that decision and codify it once, so the next builder doesn't start from scratch.
The agent deployed the pipeline. Whether an agent is ready to ship or needs to sit on the bench shouldn't be a gut call. You build the opinionated framework, and the checks every agent has to pass, that decides it.
The engineer/agent builder experience. A simple set of skills or agents that walk a builder through the whole process: which technology to reach for versus when to keep it simple, what deployment strategy a given agent needs. You make building on the harness delightful, not a scavenger hunt.
You'll work at the seam between the surface people use (Flow — one queue for every job, prospecting to claims), the operating map underneath it (Rails & Trails — the SOP as a live state machine where worn manual paths get paved into automation), and the shared platform services below that (the Substrate — data in, testing on real cases, metrics by default). You'll also partner closely with Applied Science, who are training their own models that this harness can leverage.
This is a builder-of-builders role. Your success is measured by what everyone else ships: how fast a domain engineer or a DRI can put a new agentic capability in front of an underwriter, and how confidently we can turn its autonomy up.
What you'll ship
First 30 days: Define and implement the skill-and-eval contract, running end-to-end for one real vertical (e.g., underwriting referrals or inspection verification). Ensure decision capture is live from day one, providing transparent, actionable results to the underwriters.
First 3 months: Build an opinionated platform for agent deployment and lifecycle management, providing the entire engineering team with a unified, scalable way to ship new capabilities.
First year: Develop a high-throughput orchestration engine designed to manage agents and evaluations at scale. This includes coordinating complex, multi-agent workflows, managing dependencies, automating fleet-wide health monitoring, and ensuring consistent performance across the entire agentic ecosystem.
Core competencies (Required)
Proven experience shipping agentic systems into production environments where reliability is critical.
Demonstrated track record of building tools or frameworks for engineering teams.
Proficiency in designing and managing evaluation frameworks.
Hands-on experience building Model Context Protocol (MCP) implementations.
Experience managing prompts and skills in production environments.
Familiarity with authorization frameworks and best practices.
Comfort and experience in product discovery.
Nice-to-have skills
Experience in insurance or another regulated operational domain.
Learning from logged human decisions: Transform unstructured operator event data into preference datasets for fine-tuning and reward modeling.
Built internal tooling for agentic software development — sandboxed environments, fast worktrees, anonymized production data for testing, agents that verify their own changes.
Enough frontend range (React/Next.js) to prototype the queue surface a skill feeds, rather than handing off a swagger doc
Contributions to open-source agent frameworks, eval tooling, or MCP servers.
Experience partnering with an applied-science or research team and shipping their work into a product loop.
You're probably not a fit if
Your agentic experience is primarily conversational assistants or RAG chatbots.
You want a fully specified backlog. This role is defined by ambiguity for the first year.
You'd rather build one impressive agent than the boring, reliable substrate that lets fifteen people build theirs.
Compensation
The annual base salary range for full-time employees in this position is $240,000 to $295,000 + meaningful Equity Grant.
Compensation decisions are dependent on several factors including, but not limited to, an individual’s qualifications, location where the role is to be performed, internal equity, and alignment with market data.
Benefits:
Above-market Health, Dental, and Vision coverage
Weekly lunch stipend
Flexible time off + holidays
401(k) plan
Commuter benefits
PAT & MAT Leave
Short-Term and Long-Term Disability
Monthly team gatherings
In-office perks
Work Authorization
Candidates must be authorized to work in the U.S. Stand does not sponsor new work visas. We can consider candidates on TN visas, O-1A visas, or H-1B transfers with three years or more remaining.
Equal Opportunity Employment
Stand is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. We believe that diversity enriches the workplace, and we are committed to growing our team with the most talented and passionate people from every community.
We are committed to providing reasonable accommodations for qualified individuals. If you require assistance
Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.