Work

AI Product Manager and builder based in Dubai. I shipped Vertex AI into production at Mashkor, and now most of what I build is AI agents — across conversation, research, ops automation, and decision systems — with the governance and evals that make them work in production, not just demo.

My strongest work sits at the intersection of product judgment and technical execution. I came up as a software engineer, so I build the prototypes myself: agent systems with tool-use and guardrails, recommendation systems, eval loops, and API-driven tools. I care about the full loop — user problem, model behavior, quality checks, roadmap trade-offs, launch, and measurable business impact.

t = +5.35

Risk filter validated out-of-family (2,250 gated signals of 3,486) — published with its failure conditions

3×

MAU at Mashkor · shipped Vertex AI (+15% activations)

5×

Growth at PriceLabs · team 7 → 150+

AED 2M

Built 0→1 at Dash Capital in 18 months

Product Management Growth AI & LLMs Fintech Web3 Dubai, UAE

Hiring, or want to talk AI products? WhatsApp me — I usually reply within the hour. Or ask my AI clone for the short version.

Glasshouse · Building · Shipping AI agents
Jan 2026 – Present

AI Product Manager · Dubai, UAE

I find things that hold up under adversarial testing — then publish where they break. A market-wide risk filter confirmed on a second, independent strategy family (t=+5.35 on 2,250 gated signals). I build the governance that makes agents safe to run unattended.

  • → Validated a regime filter out-of-family rather than in the data it was found in — t=+5.35 on the 2,250 gated signals (of 3,486) — and published its limits alongside it: it dies above 0.20R of friction and has decayed year over year. An edge without stated failure conditions is a story, not a finding
  • → Modelled expected value before committing engineering time: a three-year, fees-included five-mode A/B chose the production design on evidence — PF 2.73 against a 2.35 incumbent, bootstrap 90% CI [+$503, +$1,066], consistent across 8/8 instruments and every year tested
  • → Measured a signal edge at 82.8% accuracy across 25,901 markets — and say plainly that execution fillability is still unproven, because a signal result is not a tradeable one
  • → Built the evaluation-and-promotion pipeline behind all of it: 60+ candidates assessed against criteria agreed before results existed, 4 promoted — the gate is what makes the four worth believing
  • → Systematized capital allocation across 8 parallel workstreams — a centralized layer with reservations and conflict / duplicate guards that eliminated double-allocation and the drag that was quietly eating returns
  • → Operate AI agents in production with staged promotion (shadow → paper → live), tested kill-switches, human-in-the-loop, and expected-throughput watchdogs that cut silent-failure downtime to near-zero — wrote up the approach as RFC #7218 on preventing catastrophic agent actions (open-source, 8/8 tests)
  • → Claude-native daily — Claude Code (hooks, slash commands, MCP servers), the Anthropic Agent SDK, tool-use, prompt engineering, LLM evals; came up as a software engineer (Java / microservices), so I build the prototypes myself
  • → On the side: pavan.blog (Claude-powered conversational agent) · Angel portfolio: xAI, GrowthX, WorldMobile, Worldcoin
Dash Capital · Founder — Product & Operations
Aug 2024 – Dec 2025

Founder-built venture · Dubai

Founded and grew a business 0 → AED 2M in 18 months — by building the AI ops product that ran it.

  • → Built the AI-powered operations stack that ran the company — automated client onboarding (KYC/compliance-aware), buyer/seller outreach sequencing, CRM and lead-gen workflows — replacing what would normally need a 3–5 person ops team
  • → Built dubai-re-intelligence (open-source): a pipeline turning raw DLD transaction data into decision dashboards that drove every allocation call — now a live, refusal-aware Q&A demo over real 2026 sales (85% on an independent eval at first contact)
  • → Full P&L ownership from zero to AED 2M (~$545K) annual revenue in 18 months — founder-operator, not a hired role
  • → The takeaway: the part I loved most was building the systems — which is why I went all-in on AI product
Mashkor · Senior Product Manager (Growth)
Nov 2022 – May 2024

Hyperlocal · B2C · Kuwait

MAU 3× (7K → 25K), revenue 2.8× in 18 months.

  • → Led OKR strategy — roadmap focused on acquisition, retention, engagement, and revenue growth
  • → Improved activation loops by 40% in 8 months via journey mapping, onboarding, and ARIA framework
  • → Shipped ML recommendation engine (Google Vertex AI) — owned feasibility, eval-set design, ranking output iteration, A/B framework, +15% activations, +1.3× engagement
  • → Identified SOM of 500K users, developed two core personas for "Buy Anything" and "Pick Up Anything"
  • → Cross-functional leadership: UI/UX, Engineering, Marketing, Customer Support, Finance, Legal
Nova Benefits · Growth PM
Jan 2022 – Oct 2022

Insurtech · B2B · India

2.5× website traffic, +30% product leads in 4 months.

  • → Pioneered product-led website initiative — improved lead generation by 20% in 4 months
  • → Built LinkedIn ABM campaigns that amplified product leads by 30%
  • → Automated sales funnel and refined outreach — reduced response times by 25%
  • → Achieved 80% OKRs for two consecutive quarters through growth experiments and process improvements
rtCamp · Growth Specialist & PM
Aug 2020 – Dec 2021

Enterprise Web Agency · Remote

+30% template discovery for 100K+ installs in 3 months.

  • → Optimized Google Web Stories plugin using competitive search strategies
  • → SEO consulting for HCL — measurable improvements in search visibility within 3 months
  • → Upgraded digital web practices for enterprise clients — ensured effective crawling and indexing
PriceLabs · Growth & Product Consultant
2020 – 2021

Consulting · Remote

Growth & product partner through a 5× scale-up.

  • → Growth and product consulting for PriceLabs (vacation-rental AI SaaS) through its scale-up phase
  • → Company grew from 7 to 150+ people across the years around the engagement
  • → Delivered against the founders’ roadmap — client-side product work, remote
Flint Technology & Systems · Growth & Product Lead
Dec 2011 – Aug 2020

Boutique consultancy · Remote

300K users across 6 content platforms. 240% listings growth for a marketplace client.

  • → Grew listings 240% (50K → 170K) for an online property marketplace
  • → Built content websites with affiliate marketing — 300K users across 6 platforms
  • → Sold 2 content websites within a year at premium valuations
Softronikx · Director, Project Management
Apr 2008 – Nov 2011

Pune, India

Business strategy, market analysis, and revenue growth.

  • → Developed and executed business strategies including sales and business planning
  • → Identified new opportunities and drove revenue through market analysis
  • → Led business analysis to streamline operations and improve performance

Systems I've Shipped

Each system below is a product decision — what to build, what to gate, and what not to build. The domain varies, the judgment pattern is the same.

Autonomous Execution Agents

Python · Multi-strategy decision engine · Live venue execution · Staged promotion

A decision system is only as good as the discipline around what it is allowed to do. Built a multi-strategy execution layer where every candidate strategy runs the same lifecycle — shadow, then paper, then live — and is measured against pre-registered criteria it declared before seeing results. Most candidates never graduate: over 60 were evaluated, 4 promoted, and more than 90% killed by their own test batteries. Said no to promoting on a good backtest — a strategy earns capital by surviving falsification, not by looking good in-sample.

t=+5.35 · 2,250 gated signalsPF 2.73 · CI [+503, +1066]60+ evaluated → 4 promotedPre-registered criteria

Glasshouse — Transparent Research Desk

Agent infrastructure · Falsification test batteries · Non-custodial copy service

Most of this industry markets performance and hides method. Built Glasshouse on the inverse: the product is verifiability. Systematic strategies are developed under falsification discipline and killed by their own test batteries before they touch capital, and the copy-service model is non-custodial — clients keep custody of their funds and compensation is performance-share only. The whole operation runs on the agent infrastructure below: research agents, monitoring loops, promotion gates. Said no to performance marketing — the pitch is the audit trail, not the returns.

Launched 2026Non-custodial by designPerformance-share onlyRuns on own agent stack

pavan.blog — Digital Clone

Astro · Claude API (claude-opus-5) · Tool use · Prompt caching · SSE streaming · Vercel

Conversational AI clone on pavan.blog. Designed the system prompt intent-first: it works out who the visitor is (recruiter, collaborator, exploring) before answering, then pulls only the knowledge that fits that need through tool calls, and hands genuine leads to me. Every conversation logs its token cost (input, cache, output) so cost per conversation is measured, not guessed; prompt caching keeps the large stable prefix cheap. Said no to RAG — the corpus is small enough that on-demand knowledge sections are cheaper and simpler than a vector store.

Live on pavan.blogIntent-first promptLead captureToken cost logged per chat

AlphaGrid — Production Orchestration Layer

Python · Flask · systemd · Webhook signal routing · Telegram alerts

Autonomous systems that move money need a guarded layer between the decision engine and execution — otherwise a model bug becomes a wallet bug. Built and operate a Python orchestration service routing signals from upstream decision systems through a risk guardian (drawdown-kill, per-strategy loss caps, conflict and duplicate guards, two-stage entry, tested kill switch) and staged-promotion gates (shadow → paper → live). Telegram alerts on every entry, close, and error. Said no to hooking every upstream system immediately — only ones that pass the pre-production gate are enabled. Public pattern extracted to GitHub; live dashboard and production adapters remain private.

Multi-stream productionRisk guardianKill switch testedTelegram alertsStaged-promotion gates

Lab Framework — Multi-Stream Promotion Infrastructure

Python · Flask · cron · Plug-in stream registry

Once you scale beyond two production streams, ad-hoc promotion decisions become the bottleneck — and the source of every avoidable incident. Built a Lab Framework where each candidate stream registers its own gate criteria (statistical thresholds, capital limits, error tolerances), and a nightly review cron measures every stream against its criteria. Two endpoints surface the state of the world: /api/live-readiness reports which streams have passed all gates, /api/risk-status reports which need attention. Said no to manual promotion overrides — every promotion is gate-driven and audit-logged.

Stream registryNightly review cron/api/live-readiness · /api/risk-statusPlug-in patternAudit-logged promotions

Dubai RE Intelligence

Python · Pandas · DLD open-data pipeline · Eval harness · Vercel

Dubai real-estate decisions at Dash Capital were being made against scattered DLD exports — slow and hard to re-run. Started as a toolkit that loads and normalises DLD transactions; now a live Q&A over real 2026 residential sales. The hard product call was what it must NOT answer: the DLD feed only serves the current year, so year-on-year is refused rather than estimated, modelled yields are labelled as modelled, and a question it cannot answer from the data gets a refusal with answerable alternatives — never a confident guess. Evaluated on a question set written by a separate model that never saw the code: 85% (34/40) at first contact. Live questions feed a human-gated learning loop — failures become eval candidates, a person approves the label. Said no to an LLM-only answer path: a rules planner answers at $0 per question, and the LLM planner is benchmarked against it on the same evals before it earns a place.

Live demoRefusal-aware85% independent evalHuman-gated learning loop

Insight Bay — WhatsApp AI Lead-Responder

Claude API · WhatsApp Business · Booking & CRM workflow automation

A UAE field-services business was losing leads to slow replies and spending hours daily on inquiry handling. Built (as a paid engagement) a WhatsApp-native AI assistant that answers inquiries, qualifies leads, quotes and books jobs 24/7 — with human-in-the-loop handoff for edge cases and pricing exceptions. Response time went from hours to seconds, and the owner reclaimed hours of admin every week. Said no to a custom app — met customers on the channel they already use.

Paid client engagementResponse time: hours → seconds24/7 lead captureHuman-in-the-loop handoff

Content Research Agent

Python · Anthropic Claude API · Cron-friendly

Content research is the slowest, most expensive part of running a niche newsletter when done by hand. Built two Claude-powered agents that turn a Monday morning's research into a 2-minute cron job: one agent surfaces trending topics, pain points, and regulatory updates (VARA, UAE Central Bank); the other runs a YouTube content-strategy brief with hook titles and content gaps. Said no to RAG, scraping, and vector DBs — a single structured prompt is enough for weekly cadence. Keep the complexity budget for downstream production.

Claude Opus 4.7Env-var configDated reports2 agents, one pattern

How I Build — Playbooks

The patterns behind the systems above, written up in depth. Every piece comes from something running in production — not theory.

Case Study: Shipping AI Agents That Act Without a Human

Glasshouse — two years in production

How I decide whether an AI agent may act on its own — the shadow → supervised → autonomous ladder, a quality bar that killed 90%+ of my own candidates, and three products I stopped with the numbers shown. Includes the 30-hour outage where nothing errored and nothing ran. Written for product people.

The Architecture of a Self-Driving System

The whole fleet — architecture overview

Seven layers that make a fleet of agents safe to trust — proposal, validation, execution, coordination, memory, monitoring, human — held together by two rules: no agent promotes itself, and no agent grants itself resources.

Evals for Agents That Act

Lab Framework · AlphaGrid

You don't score an agent's output — you grant authority in stages. The shadow → paper → live promotion ladder, with falsification gates at every rung. This is the pipeline behind the Lab Framework and AlphaGrid.

Risk Guardian: Preventing Catastrophic Actions in Long-Running AI Agents

AlphaGrid execution guard · RFC #7218

An evaluator that catches a bad output after the fact is too late — you can't un-send a webhook. RFC #7218: a deterministic pre-action safety gate (budget caps, duplicate guards, two-stage dispatch, drift monitor, kill switch).

Building a Guardrailed AI Agent with Human-in-the-Loop

WhatsApp AI Service Assistant

The gate that decides whether an agent acts alone: self-evaluate on confidence AND sensitivity, auto-execute only when both clear, otherwise escalate to a human whose decision feeds back. The same pattern behind the WhatsApp service assistant.

What a Year of Running Production AI Agents Taught Me About Reliability

All production agents

Five hard-won lessons: agents fail silently (watchdog the absence of activity), backtest ≠ live, build the kill switch first, multi-agent needs a global off-switch, and the reasoning trail is the most valuable output.

I Built an n8n Workflow from Claude Code, via MCP

Agent tooling · MCP

What "agent-friendly interfaces" means in practice: an agent building on a real platform through MCP — where the tooling helps, where it fights you, and what platform teams should take from it.

Angel Investing

→ xAI — Elon Musk's frontier AI company
→ GrowthX — community & education for growth professionals
→ WorldMobile — decentralized telecom infrastructure
→ Worldcoin — crypto identity protocol

Education & Certifications

McKinsey Forward

McKinsey.org · 2025

Gen AI Product Strategy

Walmart AI Leaders · 2024

Gen AI — Idea to MVP

Uber AI Leaders · 2024

Product Strategy

Reforge · 2023

Master in Product Management

Reforge · 2022

Product & Growth Bootcamp

GrowthX · 2021

Advanced Google Analytics

Google

Advanced SEO Strategies

UC Davis

Content Marketing

HubSpot Academy

Education

Bachelor of Engineering (BE) in Information Technology — University of Pune, 2006

Tools & Skills

AI & Agents

AI AgentsAgentic WorkflowsClaude CodeMCP ServersAnthropic Agent SDKTool-use / Function-callingLLM EvalsAgent GovernanceGoogle Vertex AIAnthropic Claude APIOpenAI APIModel Selection (cost / latency)Prompt EngineeringRAGHuman-in-the-loopResponsible AI

Product & Growth

A/B TestingOKRsPRDsGTM StrategyPLGActivation FunnelsExperimentationUser ResearchRFM Analysis

Analytics & Tools

AmplitudeMixpanelMoEngageLookerFigmaNotionCursorSEO

Get in touch

Hiring, collaborating, or want to talk AI products? The fastest way to reach me is WhatsApp — I usually reply within the hour.