Shipping systems for quoting, logistics ops, and B2B pricing
I identify structural bottlenecks behind operational chaos and build systems that make fixes permanent. This pattern repeats across every domain I've worked in:
Engagements: Product Lead, Technical Product Owner, and AI Product Manager roles where diagnostic thinking + system building + outcome measurement are the job. Ideal for ops-AI companies, system-heavy startups, and roles that bridge technical complexity with commercial clarity.
Three themes run through this work: quoting systems (CV + OR for logistics ops), ops & energy modeling (scenario-first deployability), and B2B pricing architecture (tier consolidation, take rate, ARPU). Expandable case summaries below; separate deep-dive pages cover eval harnesses, pricing ladders, and product notes.
Manual job quoting took dispatchers 1 hour per estimate. Quotes were unreliable, causing margin erosion from inaccurate pricing and losing customers to faster competitors.
Commercial impact: Checkout abandonment was high, and operational costs scaled linearly with volume, hurting cash flows and unit economics.
Last-mile moving and delivery was still quote-by-phone. Customers called 3-5 movers and booked whoever called back first with a price: speed-to-quote was the primary booking driver, not price. Digital-first competitors (TaskRabbit-style platforms) were taking share not because they were cheaper but because they were instant. The opportunity was to out-quote, not out-price.
Focus groups with bookers surfaced that the dispatcher call was a trust barrier, as customers felt sized up and upsold before any price anchor. Session replay (Heap, Clarity) showed 60%+ checkout abandonment at the manual item-entry step; customers were typing dimensions they didn’t know and giving up. Tree tests confirmed customers expected a price in under 2 minutes (Uber/DoorDash mental model). Opportunity maps ranked quote-flow simplification as the highest-leverage fix over booking confirmation and crew communication.
CV approach chosen over a structured form, as simplifying the form wasn’t enough since dimension-entry was the primary abandonment driver. OR engine chosen over lookup-table pricing, since multi-item, multi-floor jobs had 40%+ margin error with static tables, causing the very dispatcher call the product was meant to eliminate.
V1 explicitly cut: real-time dispatcher chat overlay, customer subscription model, and automated route optimization (all deferred to preserve quote-flow focus).
R (Bayesian modeling), SQL, Computer Vision algorithms, Conversational AI frameworks, SMS API integration, Historical job data analysis, Predictive modeling, Operations research optimization
Founding PM reporting directly to CEO. Owned user research, product strategy, conversion, and retention. Led a team of 4 developers (1 direct report). Defined the OR engine architecture, drove research-to-roadmap decisions, and set the commercialization strategy that turned an internal tool into Quotely SaaS.
Segment instrumentation showed 70%+ of bookings came from returning customers, which shifted roadmap priority to lifecycle messaging and retention over new-user acquisition. Fill rate data revealed weekend supply bottlenecks, adding crew availability forecasting to the roadmap. The OR engine’s accuracy drew unplanned inbound interest from other logistics operators, and that signal drove the decision to commercialize it as Quotely rather than keep it proprietary.
AI systems succeed when they reduce operational ambiguity, not when they maximize sophistication. Structured information capture and reliable handoffs beat open-ended interaction. The OR engine’s commercial value was unplanned, emerging from measuring fill rate rather than from a product roadmap.
I designed a four-layer eval architecture for Quotely: vision, catalog matching, logistics snapshots, and end-to-end quote accuracy, with offline/online flywheels and CI gates. Quotely evals →
Separately, I designed eval metrics for conversational agents: when Cohen’s κ applies, HITL recall/precision, and fixture task-success gates (Moovez SMS agent). Agent evals →
Clean-energy planning slows down when sizing, stress-testing, and finance run in disconnected tools: assumptions drift, comparisons break, and nobody can defend a single narrative from deck to data room. During my MBA I built spreadsheet models to reason about deployability; they surfaced the questions but stayed siloed.
Engineering risk: Naively porting workbook logic to a UI yields surface parity without semantic correctness (hidden unit mismatches, unsafe handoffs between sizing and stress, economics that ignore feeder reality).
Project Epsilon connects those worlds with one version-controlled
Scenario that feeds every engine, eliminating copy-paste drift between models.
CapacitySeed from sizing into stress and
financePositioning: Planning and diligence for data centers, microgrids, and RTO-style stress, rather than a real-time EMS.
Python: custom PyPSA linear programs for 24/7 CFE and microgrid
sizing; ISO-RTO Monte Carlo stress testing; agent-based VPP dispatch;
AC power flow / distribution validation. Typed Scenario contract,
EngineRegistry, ScenarioStore, normalized finance assumptions (e.g. WACC,
battery efficiency) for cross-engine comparability.
Delivery: Public app hosted on Cloudflare Workers (epsilon.xblavania.workers.dev).
Product builder and engineer: learned distributed-energy dynamics deeply enough to shape module boundaries (what must be true for sizing, stress, finance, and network to tell one coherent story). Implemented custom PyPSA LPs, ABM-style VPP dispatch, Monte Carlo stress paths, and end-to-end architecture that bridges physical network validation with financial planning. Owned scenario contract, adapter patterns, and assurance story, rather than just a spreadsheet port.
Outputs are for analysis and professional judgment only; they do not constitute legal, regulatory, or FERC filing advice.
Each workflow step surfaces a risk class the others cannot see. Siloed software breaks narrative coherence between pitch physics and finance; deployability is a coupling problem. The product’s job is to make those couplings explicit, comparable, and reviewable, not to maximize chart polish.
Heterogeneous buyers (advisors, brokers, appraisers) with different jobs and ACV tolerance — and no public price book. Every customer got a composed deal (module mix × seats × term × negotiated discount). The wrong move would have been discounting BVX to win deals NTV should land.
NTV (also branded Capitalization 2.0 / AltBV) was a deliberate loss-leader into an adjacent terminal-value market. BVX was the margin target SKU. Upgrade was driven by in-product feature fences — spreadsheet capitalization vs full deal equilibrium — not BVX discounting. Custom deals let NTV carry the deepest discount without training discount expectation on BVX.
In-house analytics showed fewer than 20% weekly active users despite strong demo feedback — the value metric was client-ready presentation speed, not model depth. Killed fee-benchmarking and deal-database bets; pivoted products 3–5 to presentation/export; evolved custom deal templates from usage-pattern analysis over nine years.
Revenue expanded when packaging matched the value metric. Presentation modules and feature fences drove voluntary ARPU growth — a pricing architecture outcome, not a list-price hike. Loss-leader economics only work when you instrument conversion, expansion, and blended margin.
M&A financing for sub-$10M deals is an underserved niche: traditional lenders don’t understand the deal structure, and most platforms used opaque, advisor-mediated pricing requiring a sales call to get a number. VWLL’s digital-first approach was the differentiator, but the pricing UX was recreating the friction it was meant to eliminate.
Opportunity maps surfaced that churn was concentrated at the plan-selection step, meaning users were dropping before they ever experienced the product. STP analysis revealed all three real customer segments were being served by the same two plans, making the other two tiers noise rather than value. Focus groups with M&A advisors confirmed they were pre-qualifying which clients to even recommend the platform to because they didn’t trust clients to pick the right plan without a guidance call, effectively gatekeeping the funnel. Tree tests showed users couldn’t distinguish plan value in under 30 seconds.
Consolidated from 4 tiers to 3 by eliminating the two lowest-adoption plans. The decision was structural, not cosmetic, since better labels or tooltips would not fix a plan architecture that required advisor intervention to navigate. Launching with fewer tiers meant losing upsell optionality short-term, a deliberate tradeoff validated by the research.
Pricing problems are often decision architecture problems, not just monetization problems. Simplifying choices can be more valuable than optimizing price points, and advisor behavior change, not customer behavior change, was the real unlock.
K-12 content filtering and student safety monitoring is a trust-sensitive, compliance-driven market. Districts purchase under regulatory pressure, but adoption depends entirely on whether staff act on alerts. The competitive dynamic is not feature-depth, but rather whether the system generates a signal that counselors and IT admins are willing to act on. Alert fatigue had already collapsed trust at similar products district-wide; Netsweeper was at risk of the same outcome.
Focus groups with district IT admins and school counselors revealed the core behavioral problem: staff had stopped opening the alert dashboard entirely, because 8 out of 10 alerts had been false positives long enough that checking them felt like wasted time. The research surfaced a competing hypothesis from the engineering team: add new detection categories (more coverage = more value). The research showed the opposite: more detection with the same false-positive rate would accelerate trust collapse. Tree tests showed counselors couldn’t locate high-severity alerts, which were visually buried under low-severity noise.
Reframed the product’s core success metric from “detection accuracy” to “operational trust”, a shift that unlocked the correct prioritization framework and resolved the roadmap standoff with engineering. Deprioritized all new feature development for a full quarter. Redirected the outsourced team entirely to false-positive pruning, alert workflow redesign, and structured QA cycles.
Explicitly rejected: the proposal to add new detection model categories until existing signal quality was restored.
Detection systems fail when users stop trusting the signal. Accuracy metrics alone are insufficient. Operational reliability matters more than theoretical performance, and the right metric unlocked the correct prioritization framework over a roadmap standoff with engineering.
Loss-leader ladder, feature fences, and custom deal templates across a nine-year SaaS portfolio without a public price book.
Four-layer eval system for vision, catalog matching, logistics snapshots, and end-to-end quote accuracy with CI gates.
When Cohen’s κ applies, HITL recall/precision, and fixture task-success gates for SMS support agents.
BVXpress product for ICI clients: workforce ingestion, diligence reporting, and transaction compliance workflows.
Patient and clinician experiences designed for white-label deployment to a regional government health system.
Search-first discovery for physical appliances, virtual appliances, and service add-ons with Mixpanel + Pendo instrumentation.