Fit the model to the hardware you can afford.
Adapt the model and serving engine where measurements justify it, install the system on your hardware or cloud account, and hand over the configuration and operating instructions.
- Engagement
- Scoped to your workload
- Price and dates
- Agreed in your proposal
- Support
- With the researcher who builds it, under agreed terms.
Who this is for
- You have a defined workload and a hardware or operating-budget constraint.
- You need a deployment you can operate, not just a model that loads.
What you receive
- An installed deployment within the agreed scope.
- The configuration and a measurement record.
- Operating and recovery instructions.
- A handover and agreed support plan.
What we need from you
Representative requests, quality and response-time targets, hardware or account access, and the person who will operate the result.
How it runs
Measure the baseline
Establish what the current model and system do on your workload.
Change what limits the system
Test quantization, pruning, architecture and engine choices against quality as well as speed and memory.
Install and hand over
Test the deployment in its intended environment and document how to run it.
Cloud setup is an option
We install on your cards or set up the deployment in your cloud account. Nebius (UK) and Verda (Finland) are our production references; the recommendation includes compute, engineering and ongoing operation.
Read about measuring the whole request
memra is the lab’s from-scratch Rust + CUDA inference engine for Blackwell. Its kernels, quantization arithmetic and speculative decoding are developed and measured in public.
This is research provenance, not a claim that this engine runs the deployment trial or predicts your results.
Scope, handover and support
Your proposal sets out the work, price, start and handover dates, and what is needed from you. Hardware, cloud charges, engineering and ongoing operation belong in the cost discussion. On-site days and continuing support are agreed as part of the engagement. The researcher who builds the system keeps supporting it under those terms.
Questions about the work
How is the price decided?
From your workload, hardware and the implementation and support scope we agree together. Your proposal separates the engineering work from hardware, cloud and ongoing operating costs.
Who keeps supporting the system?
The researcher who built it. We agree support coverage, on-site work and how to handle problems before starting.
Do I need to buy hardware first?
No. Start with your workload and budget. We compare the options before recommending a purchase or deployment.