Company

We optimize the number the enterprise actually cares about.

VectorStackAI is a research-led team building the discipline of end-to-end GenAI optimization: tuning retrieval and agent stacks against the product metric, not the leaderboard. This page is who we are, and why the company exists.

Founded 2024 · Amsterdam
01 / Mission
“Most of the GenAI market competes within a layer. We integrate across them.”

Every layer of the GenAI stack is a market of its own: embedding models, LLMs, vector databases, rerankers, agent runtimes. Each vendor competes horizontally, benchmarked against its own layer, while the enterprise has to integrate the layers vertically into a product. The two don’t line up: as our own engagements keep showing, the horizontal winners are not guaranteed to be the right parts for a vertical build. So we build on the other side of that gap: nothing we sell is an isolated component. PreciseSearch is the whole retrieval slice (dense and sparse embeddings, index, reranker) tuned as one system; AgentGrad optimizes the whole agent harness against its end task. Vertical integration is the product, at every level we offer it.

The urgency is a product of the rush. Stitched-together APIs are enough to launch a first iteration, and right now everyone is launching. But a stack of standalone APIs cannot learn from your feedback: nothing in it adapts when production tells you an answer was wrong, slow, or expensive. The eventual winners will own and optimize their stack against that feedback. We exist to make that ownership possible, with custom components and IP that turn a stitched demo into an optimizable product.

One principle governs the work: improvements flow top-down from the product metric. That requires two things a standalone API never gives you: modules in the stack that can adapt, and optimization done directly at the operating point where you deploy.

Where this goes: an enterprise shouldn’t need five vendors and an in-house integration project to get a GenAI product to production quality. They should be able to come to one place, the way cloud computing consolidated from a dozen procurements into AWS. That is the company we’re building.

02 / Founder

A decade optimizing ML across the layers it actually runs on.

I’ve spent the last ten years on one question in different forms: why does a model that wins on paper so rarely win in production? I started in research at INRIA, building and benchmarking models in the clean isolation of an academic setting, where the dataset is fixed, the metric is agreed, and the only thing that moves is the model.

Then I went to Apple, and spent roughly four years shipping machine learning end-to-end across products. That was the education. A model doesn’t ship by itself; it ships inside a system with latency budgets, memory ceilings, and a product metric that no single component controls. The benchmark winner and the thing that actually ships are, more often than not, two different models.

A model that wins a benchmark and a model that ships are rarely the same model.

At Cerebras, as a Principal Research Scientist, I worked on hardware-aware optimization of large-scale ML: co-designing across the boundary between the model and the silicon it runs on. The same pattern showed up again, sharper: the gains live in the seams between layers, not inside any one of them. Optimize a layer in isolation and you leave most of the headroom on the table. Optimize across the boundary and the numbers move.

GenAI is that lesson at the scale of an industry. The market sells the enterprise a stack of best-in-class components (an embedding model, a vector DB, a reranker, an agent runtime), each tuned for its own layer, none of them tuned for the product. It is trivially easy to stitch them into a v0 demo. Taking that demo to v1, at the accuracy and latency a real product needs, is where most enterprise GenAI projects quietly die.

VectorStackAI exists to close the v0 → v1 gap, and to leave you owning the result.

That gap is what VectorStackAI closes. We take your KPI, optimize the whole stack against it, and hand back a Pareto frontier and the exact config that gets you there, plus the proprietary assets that make the result yours, not a recipe a competitor can copy. It’s the discipline I wish every team I’d worked on had been able to buy.

03 / How we work

We start with the constraint, not the stack.

An engagement is a search for your frontier: instrumented, end-to-end, and owned by you at the end.

01

Name the number

The first conversation is about the KPI, not the architecture. Accuracy floor, latency budget, hallucination rate, cost ceiling: we anchor on the metric your product is judged by, and the operating point you sit at today.

02

Map the frontier

In the first weeks we instrument your current stack, establish the v0 baseline, and search the configuration space: swapping components, fine-tuning where it moves the metric, and accumulating operating points on a clean accuracy-vs-cost plot.

03

Hand over the assets

You pick an operating point on the frontier. We hand back the exact config plus the proprietary assets behind it: fine-tuned embeddings, learned weights, distilled rerankers, custom indices. Owned by you, not licensed from us.

Tell us the constraint you're stuck against.

Send the KPI: accuracy, latency, cost, hallucination rate. We'll tell you where your frontier is. One line, one reply, no funnel.