Skip to content
Blog
5 min read

Case study: how a bank in Eastern Europe went from 33 to 43 tok/s

A bank in Eastern Europe faced a challenge common to financial institutions running AI in a regulated environment: it needed faster inference, without giving up control of its own data.

A BigTech, working out of a lab in Luxembourg, had taken two months of work to bring the operation to 33 tokens per second (tok/s) β€” the pace at which the model generates responses. It was a functional result, but still far from what the available hardware could deliver.

Applying heterogeneous computing β€” orchestrating the accelerators already available in the bank's infrastructure instead of relying on a single processing vector β€” MultiCortex raised performance to 43 tok/s.

That jump from 33 to 43 tok/s represents a 30% productivity gain on the very same infrastructure, with no hardware replacement and no loss of data control for the bank.

The case illustrates well what MultiCortex sets out to do: extract the maximum from every chip already installed, rather than treating more hardware purchases as the only answer to performance.