Avi Fenesh's research lab
Self-deploy frontier models on hardware you can afford.
Tiyuvta helps you choose, adapt and deploy open-source frontier models for agent systems and other workloads. We work on the model, its architecture and the serving engine to reduce cost and improve performance on your cards, or in a cloud account we set up with you.
Choose the work you need
Start with the decision or build in front of you.
Model & hardware assessment
Know what to run before you buy hardware.
Candidate comparison · Workload baseline · Cost recommendation
Model deployment & optimization
Fit the model and serving engine to hardware within your budget.
Installed configuration · Measurements · Operating instructions
Fine-tuning & evaluation
Test the quality gap, then train where the evidence supports it.
Evaluation set · Candidate comparison · Deployment artifacts
Agent-system builds
Connect models to your tools and data in a system you can run.
Integrations · Permission and approval rules · Workflow tests
What fits your hardware?
A model fitting in memory is only the start. The choice also depends on answer quality, request size, simultaneous work and the cost of operating the system.
Inputs
- Task and examples
- Quality checks
- Expected load
- Hardware and budget
Decision factors
- Model choice
- Memory and execution
- Response time
- Total operating cost
Output
- A measured recommendation and a deployment scope.
From examples to a system you can run
Choose
Define the task and compare model and hardware options.
Adapt
Test fine-tuning, architecture and engine changes where they address a measured gap.
Deploy
Install, test and document the system on your hardware or cloud account.
Support
Keep working with the researcher who built it, under agreed terms.
Research you can inspect
The lab researches engines, quantization, pruning, speculative decoding and Hebrew models. Published experiments describe their own setup and limits; they are not promises about your deployment.
memra is the lab’s from-scratch Rust + CUDA inference engine for Blackwell. Its kernels, quantization arithmetic and speculative decoding are developed and measured in public.
This is research provenance, not a claim that this engine runs the deployment trial or predicts your results.
Work directly with Avi
Avi Fenesh is the researcher and engineer behind Tiyuvta. You discuss the work with the person who builds and supports it.
Bring the workload and the budget.
We will work out what to measure and what to build.