Case study: how a bank in Eastern Europe went from 33 to 43 tok/s
A bank in Eastern Europe faced a challenge common to financial institutions running AI in a regulated environment: it needed faster inference, without giving up control of its own data.
A BigTech, working out of a lab in Luxembourg, had taken two months of work to bring the operation to 33 tokens per second (tok/s) β the pace at which the model generates responses. It was a functional result, but still far from what the available hardware could deliver.
Applying heterogeneous computing β orchestrating the accelerators already available in the bank's infrastructure instead of relying on a single processing vector β MultiCortex raised performance to 43 tok/s.
That jump from 33 to 43 tok/s represents a 30% productivity gain on the very same infrastructure, with no hardware replacement and no loss of data control for the bank.
The case illustrates well what MultiCortex sets out to do: extract the maximum from every chip already installed, rather than treating more hardware purchases as the only answer to performance.