Bank Β· Eastern Europe
33 β 43 tok/s
+0%MultiCortex exceeded both the bank's performance target and the benchmark achieved by a Big Tech company.
Context
The bank operated a fleet of GPUs from different manufacturers and required an output speed of 39 tokens/s. Big Tech specialists worked for two months and reached 33 tokens/s.
MultiCortex approach
MultiCortex developed a special build using heterogeneous computing to optimize AI execution on the existing infrastructure.
Result
43 tokens/s, exceeding the 39 tokens/s target and the Big Tech result of 33 tokens/s β a 30% productivity increase.