01Cost
60×
lower model cost
Per 1,000 retrieval answers
Published prices. Switch only after your quality checks pass.

modalis consulting
We help teams reduce AI costs, get faster answers and control where their data goes. You keep the code.






01Cost
60×
Per 1,000 retrieval answers
Published prices. Switch only after your quality checks pass.
02Speed
2.5×
Time to write 500 words
One request. Reading your prompt takes extra time.
03Data control
Yours
Your data. Your model.
Local jobs stay inside. You approve any external-model calls.
September 2026 pricing example: Claude Opus 5.5 at $4 per million input tokens and $20 per million output tokens; gpt-oss-20b on AWS Bedrock at $0.07 and $0.30. The calculation uses about 5,300 tokens in and 400 out. gpt-oss reasons before it answers, so real output can run longer. Generation times convert published single-request rates to 500 words: frontier APIs at 90 tokens/second from Artificial Analysis medians; a Mac Studio M5 Max at 120 with a 7B model at 4-bit from llama.cpp; an H100 at 228 running gpt-oss-20b from an independent vLLM benchmark. These are different models and setups, not a controlled quality comparison.
From your current AI setup to one your team can run. These figures show an example.
01 Audit
We map every model call your products and teams make: what it costs, how fast it is, how good it is, and what data it sends out.
You getA savings plan, based on your current costs.
$35,000 a month, mapped call by call
02 Agree
Quality, cost, speed and data rules, written down and signed before we change anything.
You getA signed delivery charter.
Delivery charterSupport assistant
03 Build
Most calls go to a small model on your own hardware or cloud. The hard ones still reach a frontier model. Each call carries only the context it needs.
You getA setup running in your cloud.
Every requestRouted by job
Context per call (tokens)75% less
04 Measure
We run both setups on the same real tasks and compare cost, speed and answer quality against the requirements we agreed.
You getA measured comparison.
Same 500 ticketsBefore → after
Quality checked against the agreed test set.
05 Hand over
Models, code, tests and runbooks, and a team trained to run and extend them.
You getReady for your team to run.
Estimate your savings
Fees and delivery
Before work starts, we agree what to build, how to check it and what it costs. For selected early engagements, the build fee is due when those checks pass.


modalis is our Mac app. It keeps a private, cited wiki of your files and AI sessions. Resume any project in a minute. Its models are trained on the app’s own tasks, never on your files.
It depends on the job. We test it on your real examples against the quality requirements we agree first. Work that needs a larger model keeps one.
No. Models can run in your existing cloud account, on hardware you own, or both. We recommend what fits your data rules and your volume.
The call is free. We look at the AI you run or plan to run, what it costs, and where data rules or speed hold you back. Then we say what may be worth a closer look.
The discovery call is free. We scope any deeper work before it starts. The proposal lists each fee and outside cost, and how we measure the result.
You do. We deploy in your cloud and your repositories and hand over the code, models, tests and runbooks. You keep control after the engagement ends.
We work inside your cloud and use your access controls. Some work can run on hardware you own. We do not use your data to train models for anyone else.
Keep them where they are the best choice. We find the calls that don’t need them, move those to cheaper or private models, and reduce how much text each call needs.
Every change has a written test on real examples, agreed before we build. A change ships only when it meets those requirements.
A free first conversation about the AI you run or plan to run: what it costs, how fast it is, and what data it sends out.