Your evals chose the wrong model
Model A beats Model B on benchmarks, but performs worse in production. Offline metrics don't capture real-world failure modes or distribution shifts.
System-level audits of model choice, evaluation, and inference — and how they interact with deployment. Cut latency, reduce cost, stabilize behavior in production. No fluff. Just measurements and decisions.
If your ML system behaves strangely, costs too much, or falls apart at scale — I've probably seen why. These are the recurring shapes.
Model A beats Model B on benchmarks, but performs worse in production. Offline metrics don't capture real-world failure modes or distribution shifts.
The model is fast, but end-to-end latency is terrible. Preprocessing, tokenization, or serialization overhead destroys performance.
Your fine-tuned model scores better but users complain. Objective mismatch or overfitting masked by evaluation blind spots.
Retrieval quality degrades at scale. Context windows overflow silently. Latency balloons with corpus size.
Nothing changed in code, but outputs degraded. Unversioned prompts and templates create non-reproducible behavior.
What worked in development becomes financially unsustainable in production. No one modeled the real cost structure.
Inputs, models, retrieval, prompts, evals, infra. We get the whole picture on one page before we touch anything.
One thing usually drives latency, cost, or regressions. We isolate it with measurement, not intuition.
Smallest change, largest effect. You leave with a prioritized list, not a re-architecture proposal.
Numbers before and after, on real traffic. If the change didn't move the metric, we keep going.
Rapid system-level diagnosis of your ML pipeline. Live session identifying bottlenecks, risks, and immediate wins.
Comprehensive analysis of your ML system: evals, model choices, inference paths, deployment, and observability.
Fix a specific critical issue. Design the solution, review implementation, unblock hard problems.
Davix Labs is led by Marcel Bischoff — a Principal Machine Learning Engineer with a background in mathematical physics and NSF-funded research.
Marcel focuses on real-world AI systems in production — diagnosing the ML decisions that break at scale and optimizing for performance, reliability, and cost. The work is small, focused engagements with senior teams who need a second pair of eyes on something expensive.
Thirty minutes. Live. Bring the system, leave with three things to fix.