About Vela

A managed semantic search engine, built to query vectors at scale without operating any infrastructure

A technical observation

Keyword search hits its limits the moment you need to retrieve meaning. Two close phrasings do not always share a single common term.

Embedding models solve that problem. They project text, images or signals into a vector space where proximity stands for similarity. What remains is querying those vectors quickly, and at scale.

That is Vela's job. You send your embeddings, we index them, and we answer similarity queries in milliseconds, up to a billion vectors per index. One single API, no cluster to administer.

2021 First index in production
2022 Second region opens
2024 Crossing a billion vectors per index
2026 Eight regions and hybrid filtering everywhere

One single API

Indexing, filtering and queries behind a single entry point

Managed service

No cluster to size, replicate or upgrade

Measurement

Published latency and recall, verifiable on your own datasets

Legible costs

Billing tied to vectors stored and queries served

Our engineering principles

Three commitments that guide every technical decision

01

Predictable latency

12 ms at the p99 on a billion-vector index. Latency holds under load, not just on average.

02

Measured recall

99.2 % recall against an exhaustive search. The speed-precision trade-off stays tunable index by index.

03

Legible costs

A single billing unit: vectors stored and queries served. No machines to provision in advance.

Interfaces available in your stack

How the service is built

Three architectural choices that explain how the engine behaves

Isolated indexes

STORAGE

Each index has its own resources and its own metadata schema, up to 4096 dimensions per vector.

Hybrid filtering

QUERIES

Metadata filters are applied during graph traversal, not after. Recall stays stable even under heavy filtering.

Regional replication

OPERATIONS

Data stays in the region you chose. Version upgrades are rolled out gradually, with no service interruption.

12 ms p99 latency
99.2 % Measured recall
40 Languages indexed
8 Regions available

Vela in numbers

0 billion vectors per index
0 max dimensions per vector
0 ms of latency at the p99
0 regions for deployment

Want to go further?

Book a technical session with our team, or write to us to discuss your use case.