tiyuvta / services
Services
Models tuned to your data, running on hardware you control, built with the people who tune the engine underneath.
Not another generic private chat. The AI service your business actually needs.
who this is for
Who this is for
Companies whose product runs on an LLM
Your API bill is cost of goods. It scales with your success, a vendor sets the price, and none of it accrues to you. We move that workload onto hardware you control, on models tuned to your task, and the bill becomes a fixed number you own. Teams that do this land near a third of what they were paying.
And teams putting AI to work inside the business
Startups and product teams that want real AI at work, not another chat window.
- a system of agents
- something that takes care of your users
- crisis triage
- paperwork and documents
- company knowledge that updates every day
- An internal helper is only one option.
And whoever will run it
Fine-tune, run on your hardware, or design the agent system with us. Rent a GPU server, buy one, or start on our endpoint, and move between them as the numbers change.
the work
What we offer, with examples.
fine-tune
Fine-tune
Adapt a model to a domain.
For example: coding tickets, or paperwork and documents.
Fine-tunebyoc
BYOC
Bring your own compute.
For example: one 5090, or your own server.
BYOCagent systems
Agent systems
Design and build an agent system as a product.
For example: users, crisis triage, paperwork, knowledge that updates every day.
Agent systems
how it runs
Rent it, or own it. Both work, and you can change your mind.
Start on rented GPUs
No capital, no commitment, and you can spend cloud credits you already hold. We tune the models, build the pipeline, and you are serving in weeks.
Move onto your own cards
A card or two in your own building runs the same models at a fraction of an API bill, with no per-token ceiling and room to grow into. We install and tune the serving engine on it, and prove the numbers on your hardware.
Which one fits depends on what you actually spend today. That is the first thing we measure, and it is usually not what people expect.
what it costs
The price depends on how much of this you want us to do.
A single starting-from number would tell you nothing, so here is the whole ladder instead. Every rung is real work we do today, and most people end up lower on it than they walked in expecting.
Published, per million tokens
Just use the API
The rates are on the page, and you never have to talk to us to get them. If your workload is small, this is the right answer and we will say so.
Talking to us costs nothing at any rung. There is no discovery fee and no paid scoping call.
Tens of thousands of shekels
We build it, then hand you the keys
We pick the models, tune them on your data, install the serving engine on your cards, build the pipeline around it and write the runbook. Then it is yours to run. The weights, the serving config and the documentation sit on your hardware, not ours.
Build price, plus a monthly
We build it and keep running it
The same system, with us operating it: capacity, upgrades, the engine underneath, and someone who answers when it matters. Sensible when you would rather not hire for this.
Into the hundreds of thousands
Specialist models that do only your work
One, two or three models trained to do the exact thing your business delivers, with none of the bloat of a general model you are paying to reason about everything else. Built to answer thousands of requests a second, in your own back office, where the data never leaves.
what one card actually does
One example, so this means something
One RTX 5090 in a machine, running Qwen3.8 27B on a serving engine tuned for that card. How fast it answers is measured on your card, with your traffic, and handed to you as a figure, rather than quoted here from a machine of ours.
The card and the machine around it cost ₪22,000 to ₪30,000, once, and a business writes a computer off against tax over three years. After that it keeps working, with no per-token bill on top and no ceiling on how much you use it.
That is enough for two or three agents working all day on what your business actually does: answering customers, moving paperwork, reading the documents nobody wants to read. You probably do not need the biggest model on the market for any of it.
And that is one card you can buy off a shelf. Bigger cards go further, and that is what we put under a workload that needs it. Which one is yours is a question we answer from your numbers, not from a brochure.
We will also tell you which rung you are on
If a small model does the job, you get a small model and a smaller bill. If your usage is genuinely low, we will say the API calls are the right answer and sell you nothing. If what you asked for is smaller than what you actually need, we say that too, before you buy hardware. Getting this wrong in either direction costs you more than we charge to get it right.
Owning the cards changes the arithmetic on its own. Computers depreciate at 33% a year under Israeli tax law, so a business writes the hardware down over three years and the real cost is well under the sticker. Spread that over four years of service life and compare it to what a metered API bills you every month while you grow.
The engineering judgement comes with the work. If we built your system, you have the person who tunes the engine underneath it for as long as you run it.
And if we built it with you, we stay with you: retunes, new models as better ones ship, and rebuilds at a fraction of the first price.
Thirty minutes, an engineer, and your actual workload. No deck.