A use case to validate
An idea for an assistant, document search or automation exists, but its feasibility, quality and cost still have to be proven on your real data.
Expertise
AI integration agency: we integrate language models (OpenAI, Anthropic, Mistral, self-hosted Llama) into your applications and processes, with RAG on your data, agents and multi-model routing, while keeping quality, cost and latency in check.
1// router.ts — illustration23export const routing = {4narrative: 'narrative model',5actions: 'instruction-tuned model',6search: 'RAG · vector database',7};89export const guardrails = [10'semantic cache', 'human oversight',11];
Situations
An idea for an assistant, document search or automation exists, but its feasibility, quality and cost still have to be proven on your real data.
The demo works, but answers vary, latency and the API bill keep growing, and nobody can measure quality.
Documents, catalogues or internal databases must feed an assistant without leaving the company's trusted perimeter.
Data residency or control over hosting rules out some APIs and points towards self-hosted models.
Scope of work
From feasibility study to operation: you can entrust us with the full integration or a specific component.
Use case, available data, quality criteria, cost per request and the role of human oversight, agreed before the first prompt is written.
Retrieval-augmented generation over your documents and business data: chunking, vector indexing, source citations and access control.
Conversational assistants and agents that call your APIs and business tools, with bounded and logged actions.
Model selection by intent, expected quality and cost, commercial models or self-hosted Llama, with a semantic cache for recurring requests.
Representative test sets, answer quality measurement and regression tracking whenever a model or prompt changes.
Integration into the mobile or web application, latency and cost monitoring, logging and procedures for drift.
Method
What the AI must produce, for whom, from which data and with what acceptable error rate. Cost per request and target latency are set at this stage.
A prototype evaluated on a set of real cases: answer quality, cost and latency. Next steps are decided on these measurements, not on a demo.
Integration into the product, guardrails, human oversight where needed, tests and continuous integration.
Quality, cost and latency tracked in production; prompts, routing and models adjusted as usage evolves.
Responsibilities
A typical division of responsibilities, agreed for each engagement.
| Topic | Your team | Black Tide |
|---|---|---|
| Data | You grant access to the relevant data and decide what may leave | We work within that perimeter and document the data flows |
| Quality criteria | You confirm what a good answer is | We build the test sets that measure it |
| Model choice | You weigh cost, quality and hosting | We compare the options on your real cases |
| Human oversight | You decide where human validation is required | We build it into the user journeys |
| Production release | You decide the release date | We prepare, deploy and monitor |
Selected work
A cascading LLM pipeline combining a fine-tuned narrative model and an instruction-tuned model, an intent-based router and a semantic cache that cuts API calls by 40%.
Read the Meduz caseAround the reading app, a Discord community with a virtual economy and a conversational agent connected to a vector database.
Read the Fantasy Alley caseEach case describes the work delivered, its scope and the current status of the product.
Budget and next steps
Use case, usable data, quality criteria, candidate models and target architecture. This is the basis for a quote.
Design and engineering time, data preparation and the expected level of evaluation, plus model API or hosting costs, estimated during discovery and set out in the quote.
Cost per request is measured from the prototype onwards. Model routing and caching keep it in check as usage grows.
Feasibility study only, a measured prototype, or full integration into your product.
Useful questions
It depends on the chosen model and how it is hosted. When data residency requires it, we recommend self-hosted Llama models. Data flows are documented during discovery.
There is no single answer. We compare OpenAI, Anthropic, Mistral or a self-hosted model on your real cases for quality, cost and latency, and can combine several.
Rarely as a first step. RAG and good routing cover most needs. Fine-tuning is worthwhile for a very specific style or task, like Meduz's narrative model.
Yes. We first examine the application, its data and its constraints, then integrate AI in increments without weakening what exists, as a project or embedded in your team.
Next step
What the AI should do, the data available and your hosting constraints: the starting point for useful discovery.