Synthetic data · RL
Dynamiks.AI
Simulated sales pipelines for training a reinforcement-learning agent, and a study of how to fingerprint them.
Reinforcement learningSynthetic dataGaussian processesMarkov modelsGenetic algorithms
The project
Dynamiks AI builds an enterprise decision system that measures and improves a company’s sales momentum. At its core is an AI agent nicknamed “the Quarterback”, trained with deep reinforcement learning to score sales opportunities, predict which deals will close, and suggest what the team should do next. It plugs into CRMs like HubSpot and Salesforce.
The catch: you can’t let an untrained agent learn on a customer’s live deals. So the agent first trains inside a CRM simulator, a “flight simulator for sales” that generates fake companies, deals, calls and emails that behave like the real customer’s pipeline. The closer the simulation is to reality, the smaller the sim-to-real gap.
What I worked on
Software Engineer Intern in San Francisco, May to August 2025.
1M+labelled synthetic CRM trajectories generated from a small real dataset
15dimensions in the trajectory fingerprint
NumPyvectorized sampling instead of nested Python loops
- The synthetic-data engine behind the RL pipeline: adaptive Markov transition models paired with Gaussian-process feature drift.
- The fingerprint study: testing whether a compact “fingerprint” of a sales pipeline is reliable enough to tell environments apart and to rebuild a matching simulation from it.
- Validation: checking the 15-dimensional trajectory fingerprints for redundancy and fidelity against held-out data before they reached the policy.
How the training pipeline works
- Fingerprint the real CRM. A customer’s real pipeline is summarized as a fingerprint: number of stages, overall conversion rate and its variance, average sales-cycle length, and how efficient the first stages are.
- Build a baseline. A “reference” environment runs the simulator with default settings, as a measuring stick.
- Calibrate. Comparing the real fingerprint with the reference gives behavior multipliers: how much faster deals move, how often they’re won, how active the sales team is.
- Simulate. A synthetic twin of the pipeline generates new opportunities and activities hour by hour, following the same patterns with entirely fictional data.
- Train in phases. The agent trains on synthetic snapshots first, then fine-tunes on the customer’s historical deals, then goes live scoring real opportunities.
The fingerprint study
Everything above depends on one question: can you trust the fingerprint? If two very different pipelines produce similar fingerprints, or the same pipeline gives a different fingerprint every time, the simulator gets calibrated to the wrong thing. I designed three tests to find out.
1. Convergence and divergence
I simulated batches of environments from five contrasting setups (standard baseline, aggressive high-velocity, conservative long-cycle, high-activity/low-conversion, low-activity/high-conversion) and tracked each fingerprint metric as the simulations ran.
- Convergence: environments with the same configuration should end up with similar fingerprints.
- Divergence: different configurations should end up clearly apart.
- Measured with spread and “closeness” statistics for every metric.
2. Bijectivity: can you work backwards?
If the fingerprint really captures a pipeline, you should be able to go from a fingerprint back to the settings that produced it. I tested that with a genetic algorithm:
- Take a target fingerprint from each of the five setups.
- Create 20 random guesses of the hidden parameters (lead-generation performance, progression rate, failure rate, average stage duration, activities per stage).
- For each guess, spin up a temporary CRM, simulate 100 opportunities, compute its fingerprint, and score how close it is to the target.
- Keep the best guesses and evolve them over several generations.
- Compare the winning guess to the true parameters, and chart how the error falls over generations and with longer simulations.
3. A confidence score for real data
Real CRM data is messy, so I built a confidence calculator that decides whether a fingerprint can be trusted before it is used for matching:
93%Stability (50% of the score): does the fingerprint stay steady as data comes in?
97%Richness (30%): how complete are the records?
89%Consistency (20%): does the data agree with itself?
Those are the final results, which works out to a weighted confidence of about 93%. High-confidence fingerprints are used for matching; low-confidence ones get flagged for more data or a manual review, along with the data-quality issues behind them.
Realistic synthetic journeys
- Markov models drive how customers and deals move between states (for example active → at-risk → churned), so transitions follow believable probabilities instead of random jumps.
- Gaussian processes generate smooth, noisy trends over time, like spending or engagement slowly decaying, instead of flat or jagged numbers.
- Together they preserve the time structure and cause-and-effect between actions and outcomes, which is exactly what an RL agent needs to learn from.
Built with
PythonNumPy / SciPyGaussian processesMarkov chainsGenetic algorithmsOpenAI GymPostgreSQLDocker