
A bank-of-the-future concept: ten features in one three.js world, voice payments checked against a voiceprint, money in chats, a shared dream jar, flight-aware cards and a filmed scene to replay with your own voice.
Overview
VOXA is a fictional bank of the future. We built its site on 28 September 2026 as a concept for banks and fintech products, in the genre of interactive feature stories where a preloader counts to ten and a camera flies through one 3D world, a feature per stop — and rebuilt everything inside that format: the models, the screens, the films and the features themselves.
The brief we set ourselves: a bank you talk to and spend time in. A voice replaces the card number, friends replace the list of payees, and a trip, a dream or a split bill becomes something people do together — the bank as your second place online after the messenger.
10
features in one three.js world, each with a working phone demo
Source: Site source files, 28 Sep 2026
56
test voice commands understood, 28 in English and 28 in Russian
Source: Parser run, 28 Sep 2026
60 fps
median frame rate resting and during flights, 1600 × 1000, Retina
Source: Measured 28 Sep 2026, MacBook (Apple M5)
0.24 MB
JavaScript gzipped, three.js included
Source: Measured 28 Sep 2026, build files
Approach
The world is one three.js scene with ten dioramas in a pastel studio whose colour flows from stop to stop; the camera flies between them along an arc, with a slight lens breathing, and follows the cursor. Every model is built in code: phones with live screens drawn on canvas, a white plane circling a boarding pass, a glass jar filling with coins, metal, glass and holographic cards, rising hearts, an arch carrying the members of an investment circle. The menu shows thumbnails rendered from the same scene when the page loads.
The voice works for real in the browser. Speech is turned into text by the browser’s own recognition service and parsed into an intent in English or Russian — transfer, request, message, dream deposit, flight, card, split, coach question — including Russian word forms such as «Ане» and «Максу». Voice ID records three phrases and builds a voiceprint from the average spectral shape and pitch, computed in the tab; later commands are checked against it before the money “moves”, and a mismatch asks for Face ID instead. The whole site can be steered by voice too: “next”, “back”, “menu”, “show the dream”.
Each feature opens a working phone app on one shared state, so money sent in one window appears in the chat, the balance and the test lab. Friends reply and react, the dream jar fills while you watch, a bought ticket becomes a boarding pass with a live countdown. The closing scene cuts two short films generated with Veo 3.1 — Alex asks VOXA to pay Anna for pizza, Anna laughs when it lands — together with the two phones on the page, then hands the microphone to the visitor. Everything is in English and Russian.
Gallery
Results
The concept is one page with ten features, ten phone demos, a test lab and the closing scene, in English and Russian. A parser run on 28 September 2026 understood all 56 test commands, 28 in each language. On the same day we counted frames with requestAnimationFrame for three seconds, twice resting on a slide and twice during a camera flight, in headless Chrome with ANGLE Metal on a MacBook with an Apple M5 at 1600 × 1000 on a Retina scale: 60.1, 59.8, 60.0 and 58.1 fps.
The page loads 0.24 MB of JavaScript with gzip, three.js included. The whole build is 4.3 MB, 3.3 MB of which are the four film clips (English and Russian versions of each).
What this does not prove
VOXA is a concept VITON13 built to show what we can make for banks and fintech products, not a commissioned project. The bank, the friends, the balances and the flights are fictional and no money moves; the two people in the films were generated with Veo 3.1 and Nano Banana Pro and are not real customers. The site has no client, users or enquiries.
Voice ID is a demonstration, not biometric security: it compares spectral averages, can confuse similar voices and was checked in headless Chrome with a synthetic microphone signal, not with a panel of real speakers. Speech-to-text needs Chrome, Edge or Safari; in Firefox visitors tap or type the commands. The frame rate was measured on one machine, and this case has no Lighthouse or SEO review.
Planning something similar? Our founder, Tarasov Vitalii, reads every message himself.
Order a similar project on VITON13 Studio
Next project
ROKOSatirical WebGL product launch site
Concept2026
All work (32)





