Salotto
A room where a group speaks Italian, with bilingual captions running underneath so a missing word does not end the sentence.
Live The room is live at salottolang.com. The first scheduled session has not been held yet.
Speaking is the thing every learner is working towards and the thing most of them avoid for years. Salotto is an hour of spoken Italian with the captions arriving as people talk, Italian large and your own language underneath, so losing a word costs you a second instead of the conversation.
What it does
Salotto is a room you join to talk. People meet in the evening, speak Italian for an hour, and the captions arrive as they talk with the Italian set large and the other language underneath. The captions are the safety net. When a word drops out from under you mid-sentence you can keep going, because the line you were reaching for is already on the screen in both languages.
The transcript is the reason the room exists. A conversation is the richest measurement of what a learner can produce, far richer than a multiple-choice answer, because the words you reach for out loud under time pressure are words you actually hold. The places where you stall and switch back to English are evidence of the opposite kind, and just as useful. Both go onto the same known-word list that Paroletta writes to and Filotto reads from before it decides what you are ready to read.
Versions are called Marks
Each Mark changes one box and nothing else: public/recognizer.js, the file that decides how speech becomes text. Mark I uses the browser’s own speech recognition and its on-device translator, which costs nothing and works on Chrome and Edge. Mark II streams audio to a vendor instead, which costs a few dollars an hour and works on every browser including phones. Holding everything else constant is what makes the comparison mean anything: when two sessions differ, the recognizer is the only thing that changed.
Who it is for
For learners who can follow spoken Italian and stall when it is their turn. It was built for a private A1 to B2 group, which is the range it suits. It is not a class, nobody is teaching, and there is no tutor to correct you.
The details
| Price | Free. |
|---|---|
| Where it runs | A browser. Chrome or Edge on a desktop for the captions to pick up your voice; Safari and Firefox can join, speak and read. |
| Account | None. The invitation link is the room. |
| Camera | Off. You join with your voice, and you are the only one who can turn a camera on. |
| Audio | Leaves your machine only to be heard by the others. The recording of your own voice stays local. |
| Kit | Headphones, which the room asks for rather than suggests. |
Questions
Do I need to install anything?
No. Open the link in a browser and you are in the room. There is no account and no app.
Why are headphones required?
Without them your machine hears the other voices through your speakers as well as your own through the microphone, and the captions stop being able to tell who said what.
Is the session recorded?
The recording of your voice stays on your own machine, and audio leaves it only so the other people in the room can hear you. Afterwards you get back what you yourself said.
Do I have to be on camera?
No. The camera starts off and stays off unless you turn it on.
What are the Marks?
Versions. Each Mark changes exactly one file, the speech recognizer, and leaves the room, the wire protocol and the transcript format alone, so two sessions can be compared without arguing about what else was different.
Get Salotto
Log: Salotto
1 entry- 22 Aug 2026 note One known-word list per person
Get updates
Occasional notes on what's happening at Enginery. No spam, no marketing.