Translate capability into product.
I work closely with engineering on model behavior and platform constraints, then turn emerging capabilities into interaction models, requirements, prototypes and product direction.
I’m Shantanu Phadke, a Staff Product Manager and former Staff Software Engineer working on enterprise AI and agentic product experiences. I’ve shipped products at scale across ServiceNow and GitHub; outside of work, I’m building a deeper technical practice around self-improving agents, reasoning, evaluation and AI-native interaction models.
I’m drawn to problems where model capability, product judgment, technical architecture and real-world deployment have to come together.
I work closely with engineering on model behavior and platform constraints, then turn emerging capabilities into interaction models, requirements, prototypes and product direction.
I came into product after progressing to Staff Software Engineer, so I’m comfortable reasoning across architecture, APIs, evaluation, data, integrations and failure modes.
My experience spans enterprise software, developer products and AI experiences where adoption, workflow quality and measurable outcomes matter.
Outside of work, I’m developing a hands-on practice around self-improving agents: reading the research, implementing the mechanisms, evaluating them and documenting what I learn.
My experience spans agentic product experiences, developer-focused AI, enterprise platforms, customer validation and product work tied to measurable outcomes.
Defining the core interaction model and roadmap for an enterprise AI agent that plans, implements, tests and iterates on application-development work.
Shaping collaborative workflows across developers, architects, reviewers and administrators, including shared context, approvals, permissions, progress visibility and human intervention.
Leading platform and extensibility strategy across skills, organizational rules, MCP-based tools, workflow data, context, memory and long-running execution.
Shipped a GenAI billing and licensing assistant integrated into Slack for roughly 400 Account Executives.
Built benchmarking and experimentation infrastructure across user feedback, sentiment, time-to-resolution and iterative product changes.
Built enterprise workflow, licensing and asset-management products while increasingly owning roadmap, partnerships, go-to-market and cross-functional execution.
Outside of work, I use public research as a starting point for hands-on experiments in reasoning, verification, planning, memory and evaluation.
Read the mechanism, implement a minimal version, instrument its behavior, compare alternatives and document the product questions that emerge.
Currently I’m working my way through Stanford CS329A, learning more about test-time compute and verification, feedback and tool use, multi-step reasoning and planning, reinforcement learning, search and open-ended self-improvement, agent memory, multimodal interaction, and long-horizon evaluation. I’m implementing the ideas that interest me most and publishing the resulting experiments, evaluations and notes here as I go.
Compare repeated sampling, best-of-N and adaptive inference budgets on a small reasoning benchmark.
Generate multiple candidates, score them with different verifier strategies and inspect confident verifier failures.
Build a small action loop that executes tools, observes feedback and changes its next move.
Compare direct generation, decomposition, adaptive branching and tree search on multi-step tasks.
Add episodic memory across repeated tasks and measure when retrieval helps or hurts.
Track success, cost, intervention and failure modes on multi-step tasks requiring recovery.
A simple loop I use to move from research ideas to stronger technical and product intuition: understand the mechanism, measure it, connect patterns across experiments, and apply the strongest insights in larger builds.
Implement small, transparent versions of mechanisms across inference, verification, planning, tools and memory.
Compare approaches using shared evaluations, cost, latency and failure analysis.
Look for recurring patterns across experiments in trust, control, observability, memory and collaboration.
Turn the strongest insights into larger prototypes and useful open-source systems.
A lightweight record of the experiments, questions and results I’m actively working through.
I’m setting up a shared evaluation harness and implementing simple inference-scaling approaches so later experiments can be compared against the same tasks, scoring, cost and latency metrics.
View current experiments ↑How much does Best-of-N help? How strong does a verifier need to be before additional sampling becomes worthwhile? How quickly do quality gains flatten relative to inference cost?
Once the baseline is stable, I’ll compare different verifier strategies and inspect the cases where a verifier is confidently wrong.
As patterns emerge from the technical work, I’m especially interested in applying them to a few broader AI product problems.
How should teams test conversational agents against hundreds of realistic users before deploying them?
What should an AI research experience look like when claims, uncertainty and contradictory evidence remain visible?
How should models work alongside deterministic systems when transforming unreliable real-world data?
What changes when an AI participates in a group decision rather than a one-to-one conversation?
What does it take to deploy an AI system into a real workflow and demonstrate measurable value?
Which interaction models become possible when the product is designed around model capability from the beginning?
A few working principles that guide how I move from model capability to product decisions, prototypes and measurable systems.
Start by asking how a new model behavior should change the user experience—not where another AI button belongs.
Use prototypes and small technical POCs to learn faster than roadmap discussion alone.
Make cost, latency, trajectories, quality and failure modes observable enough to challenge assumptions.
Evaluation, review, recovery and human intervention are product behaviors, not cleanup work.
Model novelty matters only when it changes a workflow, removes friction or creates genuinely new capability.
I prefer public research, transparent implementations and direct measurement over treating frameworks as black boxes.
I write periodically about the ideas I’m studying, the products I’m building and the questions that emerge along the way.
A running set of notes connecting current research on reasoning, verification, planning, memory and evaluation to the product questions I find most interesting.
A hybrid background that lets me move comfortably between technical systems, product strategy and business outcomes.
I’m especially interested in agentic systems, developer products, AI-native workflows and teams turning frontier-model capability into reliable products with real users.