AI Product · Agentic Systems · Builder

I build products at the edge of what AI can do.

I’m Shantanu Phadke, a Staff Product Manager and former Staff Software Engineer working on enterprise AI and agentic product experiences. I’ve shipped products at scale across ServiceNow and GitHub; outside of work, I’m building a deeper technical practice around self-improving agents, reasoning, evaluation and AI-native interaction models.

Technical PM
Former Staff SWE; comfortable going from interaction model to system constraints.
Agentic AI
Current product work spans agents, tools, context, permissions, evaluation and long-running execution.
Product at scale
Enterprise AI, developer workflows, platform products and measurable commercial outcomes.
Builder mindset
Independent research-to-code track focused on learning frontier mechanisms by implementing them.
Where I do my best work

Product problems at the intersection.

I’m drawn to problems where model capability, product judgment, technical architecture and real-world deployment have to come together.

01

Translate capability into product.

I work closely with engineering on model behavior and platform constraints, then turn emerging capabilities into interaction models, requirements, prototypes and product direction.

02

Operate at a technical level.

I came into product after progressing to Staff Software Engineer, so I’m comfortable reasoning across architecture, APIs, evaluation, data, integrations and failure modes.

03

Ship for real users.

My experience spans enterprise software, developer products and AI experiences where adoption, workflow quality and measurable outcomes matter.

04

Keep building at the frontier.

Outside of work, I’m developing a hands-on practice around self-improving agents: reading the research, implementing the mechanisms, evaluating them and documenting what I learn.

Selected product work

AI products, platforms, and enterprise scale.

My experience spans agentic product experiences, developer-focused AI, enterprise platforms, customer validation and product work tied to measurable outcomes.

SERVICENOW · 2026 — PRESENT

Staff Product Manager · Build Agent & AI Platform Core

Defining the core interaction model and roadmap for an enterprise AI agent that plans, implements, tests and iterates on application-development work.

Shaping collaborative workflows across developers, architects, reviewers and administrators, including shared context, approvals, permissions, progress visibility and human intervention.

Leading platform and extensibility strategy across skills, organizational rules, MCP-based tools, workflow data, context, memory and long-running execution.

Agentic product UXDeveloper workflowsEnterprise AIEvalsPlatform strategy
GITHUB · 2025 — 2026

AI Product Manager

Shipped a GenAI billing and licensing assistant integrated into Slack for roughly 400 Account Executives.

Built benchmarking and experimentation infrastructure across user feedback, sentiment, time-to-resolution and iterative product changes.

~400Account Executives served
20%Resolution-time improvement
$6MAnnual revenue recovered from fraud strategy
GenAI productExperimentationEnterprise adoption
SERVICENOW · 2019 — 2023

Software Engineer → Staff Software Engineer

Built enterprise workflow, licensing and asset-management products while increasingly owning roadmap, partnerships, go-to-market and cross-functional execution.

$260M new annual contract value$100M ARR tied to hardware mobile app75% setup-time reduction from NLP prototype
Independent AI work

Research, implemented and tested.

Outside of work, I use public research as a starting point for hands-on experiments in reasoning, verification, planning, memory and evaluation.

A simple research-to-build loop.

Read the mechanism, implement a minimal version, instrument its behavior, compare alternatives and document the product questions that emerge.

ReadPrimary papers + mechanism notes
ImplementMinimal working system with transparent assumptions
InstrumentCost, latency, trajectories, scores and failure modes
EvaluateCompare variants and record both positive and negative results
ExtendUse the strongest insights to shape deeper experiments and prototypes
Currently building

Learning self-improving agents by building them.

Currently I’m working my way through Stanford CS329A, learning more about test-time compute and verification, feedback and tool use, multi-step reasoning and planning, reinforcement learning, search and open-ended self-improvement, agent memory, multimodal interaction, and long-horizon evaluation. I’m implementing the ideas that interest me most and publishing the resulting experiments, evaluations and notes here as I go.

Current learning track

Stanford CS329A · Self-Improving AI Agents

Working through topics including test-time compute, verification, feedback and tool use, multi-step planning, search, memory and long-horizon agent evaluation.

Public course page ↗
What each experiment includes

A public implementation, a benchmark or evaluation, a useful visualization, concise notes on results and a clear product question.

Experiment 01In progress

Test-Time Compute Playground

Compare repeated sampling, best-of-N and adaptive inference budgets on a small reasoning benchmark.

Product question → When is extra inference actually worth the latency and cost?
Experiment 02Up next

Verifier Lab

Generate multiple candidates, score them with different verifier strategies and inspect confident verifier failures.

Product question → How should an AI product represent uncertainty when generation outpaces verification?
Experiment 03Upcoming

Feedback + Tool-Use Agent

Build a small action loop that executes tools, observes feedback and changes its next move.

Product question → When should feedback trigger retry, reflection or escalation?
Experiment 04Upcoming

Planning Search Sandbox

Compare direct generation, decomposition, adaptive branching and tree search on multi-step tasks.

Product question → Which parts of agent planning are useful for humans to inspect?
Experiment 05Upcoming

Memory-Augmented Agent

Add episodic memory across repeated tasks and measure when retrieval helps or hurts.

Product question → What should an agent remember, and what should it deliberately forget?
Experiment 06Upcoming

Long-Horizon Eval Harness

Track success, cost, intervention and failure modes on multi-step tasks requiring recovery.

Product question → What actually tells us whether an agent is becoming more capable?
Working method

From mechanisms to applied systems.

A simple loop I use to move from research ideas to stronger technical and product intuition: understand the mechanism, measure it, connect patterns across experiments, and apply the strongest insights in larger builds.

01

Understand

Implement small, transparent versions of mechanisms across inference, verification, planning, tools and memory.

02

Measure

Compare approaches using shared evaluations, cost, latency and failure analysis.

03

Connect

Look for recurring patterns across experiments in trust, control, observability, memory and collaboration.

04

Apply

Turn the strongest insights into larger prototypes and useful open-source systems.

Latest

What I’m working on now.

A lightweight record of the experiments, questions and results I’m actively working through.

Starting with test-time compute

I’m setting up a shared evaluation harness and implementing simple inference-scaling approaches so later experiments can be compared against the same tasks, scoring, cost and latency metrics.

View current experiments ↑

What I’m testing first

How much does Best-of-N help? How strong does a verifier need to be before additional sampling becomes worthwhile? How quickly do quality gains flatten relative to inference cost?

Verification

Once the baseline is stable, I’ll compare different verifier strategies and inspect the cases where a verifier is confidently wrong.

Beyond the experiments

Questions I want to explore in larger builds.

As patterns emerge from the technical work, I’m especially interested in applying them to a few broader AI product problems.

01

Voice AI evaluation

How should teams test conversational agents against hundreds of realistic users before deploying them?

02

Evidence-first research

What should an AI research experience look like when claims, uncertainty and contradictory evidence remain visible?

03

AI + messy enterprise data

How should models work alongside deterministic systems when transforming unreliable real-world data?

04

Multiplayer AI

What changes when an AI participates in a group decision rather than a one-to-one conversation?

05

Applied vertical AI

What does it take to deploy an AI system into a real workflow and demonstrate measurable value?

06

AI-native interfaces

Which interaction models become possible when the product is designed around model capability from the beginning?

How I work

Principles I use when building AI products.

A few working principles that guide how I move from model capability to product decisions, prototypes and measurable systems.

01

Capability → interaction.

Start by asking how a new model behavior should change the user experience—not where another AI button belongs.

02

Build to reduce ambiguity.

Use prototypes and small technical POCs to learn faster than roadmap discussion alone.

03

Instrument the system.

Make cost, latency, trajectories, quality and failure modes observable enough to challenge assumptions.

04

Design for failure.

Evaluation, review, recovery and human intervention are product behaviors, not cleanup work.

05

Stay close to users.

Model novelty matters only when it changes a workflow, removes friction or creates genuinely new capability.

06

Learn from the mechanism.

I prefer public research, transparent implementations and direct measurement over treating frameworks as black boxes.

Writing

Notes on AI products and agentic systems.

I write periodically about the ideas I’m studying, the products I’m building and the questions that emerge along the way.

Building my way through self-improving agents

A running set of notes connecting current research on reasoning, verification, planning, memory and evaluation to the product questions I find most interesting.

StatusFirst notes coming soon
FocusSelf-improving agents + AI product design
FormatShort essays, build notes and experiment write-ups
Education

Engineering depth + product training.

A hybrid background that lets me move comfortably between technical systems, product strategy and business outcomes.

Kellogg School of Management
MBA · Artificial Intelligence
Northwestern University · 2024
Cornell Tech
Masters · Computer Science
Cornell University · 2019
UC Berkeley
Business Administration + Computer Science
2018
Frontier AI product

Interested in hard product problems where model capability is still ahead of the interface.

I’m especially interested in agentic systems, developer products, AI-native workflows and teams turning frontier-model capability into reliable products with real users.