modalis consulting

We help teams reduce AI costs, get faster answers and control where their data goes. You keep the code.

Experience behind the work.

Delivered within
  • Nike
  • Mr. Cooper
  • Levi's
Trained professionals from
  • Salesforce
  • Meta
  • IBM
  • Microsoft

What a smaller model
can change.

01Cost

60×

lower model cost

Published prices. Switch only after your quality checks pass.

02Speed

2.5×

faster generation

One request. Reading your prompt takes extra time.

03Data control

Yours

Choose where it runs.

Your data. Your model.

  • Your cloud
  • Rented GPUs
  • Your Macs

Local jobs stay inside. You approve any external-model calls.

Sources

September 2026 pricing example: Claude Opus 5.5 at $4 per million input tokens and $20 per million output tokens; gpt-oss-20b on AWS Bedrock at $0.07 and $0.30. The calculation uses about 5,300 tokens in and 400 out. gpt-oss reasons before it answers, so real output can run longer. Generation times convert published single-request rates to 500 words: frontier APIs at 90 tokens/second from Artificial Analysis medians; a Mac Studio M5 Max at 120 with a 7B model at 4-bit from llama.cpp; an H100 at 228 running gpt-oss-20b from an independent vLLM benchmark. These are different models and setups, not a controlled quality comparison.

How an engagement works.

From your current AI setup to one your team can run. These figures show an example.

01 Audit

Find where the money goes.

We map every model call your products and teams make: what it costs, how fast it is, how good it is, and what data it sends out.

You getA savings plan, based on your current costs.

$18.4kSupport assistant$9.2kDocuments$6.1kSearch$1.3kDrafts

$35,000 a month, mapped call by call

02 Agree

Agree the requirements.

Quality, cost, speed and data rules, written down and signed before we change anything.

You getA signed delivery charter.

Delivery charterSupport assistant

Quality
Today’s answers on 500 real tickets
Cost
Under one cent a ticket
Speed
Under two seconds
Data
Stays in your cloud
Signed before work starts
Agreed

03 Build

Send each job to the right model.

Most calls go to a small model on your own hardware or cloud. The hard ones still reach a frontier model. Each call carries only the context it needs.

You getA setup running in your cloud.

Every requestRouted by job

82%Small modelIn your cloud
18%Frontier modelFor harder jobs

Context per call (tokens)75% less

Before12,000After3,000

04 Measure

Compare before you switch.

We run both setups on the same real tasks and compare cost, speed and answer quality against the requirements we agreed.

You getA measured comparison.

Same 500 ticketsBefore → after

Cost / 1,000 tickets
$41$9
Answer time
3.8 s1.1 s
Answer quality
91%92%

Quality checked against the agreed test set.

05 Hand over

The system is yours.

Models, code, tests and runbooks, and a team trained to run and extend them.

You getReady for your team to run.

Your handover4 parts
  1. 01
    Models & weightsIn your cloud account
  2. 02
    Code & servingIn your repository
  3. 03
    Tests & runbooksTo check and maintain it
  4. 04
    Team trainingTo run and extend it
Estimate your savings

Estimate your savings

What could your AI cost?

$13,560saved a month · $162,720 a year · 68% of today’s bill
check it with us A rough estimate from your numbers, not a quote. It assumes a smaller model costs about a tenth as much per call. The audit measures your real ones.

Fees and delivery

Know what you’re paying for.

Before work starts, we agree what to build, how to check it and what it costs. For selected early engagements, the build fee is due when those checks pass.

What we check
Answer quality, cost, speed and your data requirements, using agreed tasks.
What you pay
The proposal lists our fees and any cloud or model costs.
Any fee tied to savings
Agreed separately, with a limit and a defined period to measure the results.

The people you meet build the system.

our own product

Pick up a project where you left off.

modalis is our Mac app. It keeps a private, cited wiki of your files and AI sessions. Resume any project in a minute. Its models are trained on the app’s own tasks, never on your files.

whatsourced project summariesrunslocally, on your Macreleasetester build
Before you book

What buyers ask first.

01

Will a smaller model be good enough?

It depends on the job. We test it on your real examples against the quality requirements we agree first. Work that needs a larger model keeps one.

02

Do we need our own hardware?

No. Models can run in your existing cloud account, on hardware you own, or both. We recommend what fits your data rules and your volume.

03

What happens on the 25-⁠minute call?

The call is free. We look at the AI you run or plan to run, what it costs, and where data rules or speed hold you back. Then we say what may be worth a closer look.

04

What does it cost?

The discovery call is free. We scope any deeper work before it starts. The proposal lists each fee and outside cost, and how we measure the result.

05

Who owns the code, models, and data?

You do. We deploy in your cloud and your repositories and hand over the code, models, tests and runbooks. You keep control after the engagement ends.

06

How do you handle our data?

We work inside your cloud and use your access controls. Some work can run on hardware you own. We do not use your data to train models for anyone else.

07

We already use ChatGPT or Claude. Why do we need you?

Keep them where they are the best choice. We find the calls that don’t need them, move those to cheaper or private models, and reduce how much text each call needs.

08

What if quality drops?

Every change has a written test on real examples, agreed before we build. A change ships only when it meets those requirements.

Next step

Find out what your AI could cost.

Free discovery call
Twenty-five minutes on your AI spend.

A free first conversation about the AI you run or plan to run: what it costs, how fast it is, and what data it sends out.

what we discuss
  • 01your current AI use and monthly spend
  • 02where cost, speed, or data rules hold you back
  • 03what could move to a smaller or local model
25 minutes
Google Meet
America/Toronto
Choose a time

Live availability · free 25-minute call