Does the AI score how I speak, not just what I say?
Updated:
In short
Not directly. In the SwissJobs.app practice interview, the AI voice listens to you live, but the score comes afterwards and is computed from a written transcript of your answer. A second model reads that text and gives it 0 to 10, based on substance, structure and fit for the person who asked. Tone, pace, volume, accent and pauses are not measured. Delivery still matters indirectly: false starts, repetition, a missing conclusion, or an answer cut off by the three-minute limit all show up in the text and pull the structure mark down.
Practise the interview out loud, and see how you scored
An AI panel asks the questions and you answer by voice, the way the real call goes. Afterwards you get a pass probability and the moments that cost you.
Start a live practice interviewThat limit is by design, not something buried in small print. The score tells you whether your answer holds up as content. It does not judge whether you came across as calm, confident and easy to follow, and in a real interview in Zurich, Geneva or Basel that counts too.
For delivery you need a different mirror: a phone video of yourself, a colleague who will be honest, or a coach when the stakes are high. The score tells you what to say; people tell you how it lands.
- The score is computed from the transcript of your answer, not from the audio.
- It rates substance, structure and fit for the role that asked, combined into 0 to 10 per answer.
- Nothing measures tone, pace, volume, accent or pauses.
- Delivery counts indirectly: false starts, repetition and a missing ending are visible in the text.
- Mumbling or a noisy room can produce wrongly transcribed words, and the rating reads those words.
- For presence and voice, use a video recording, an honest listener or a coach.
Two steps, two different models
The practice interview has two parts that are easy to mix up. In the first, an AI voice talks with you. It plays the panel you set up at the start, asks one question at a time, and may ask one short follow-up when your answer stayed vague. This part genuinely listens: your microphone sends audio and the voice responds to what you said.
During the conversation the voice never evaluates you out loud. No praise, no corrections, no tips. That is deliberate, because a real panel does not grade your answers halfway through the meeting either.
The second part starts when you finish an answer. The transcript of the question and your reply goes to a second model that reads text only. It knows the job title, the company size and who on the panel asked, and it returns a mark from 0 to 10, two or three sentences of reasoning, and one sentence on what a strong answer would have added. When you end the session, another pass reads every transcript alongside the individual marks and estimates your chance of getting through this round.
What the transcript keeps and what it drops
The transcript holds the words you said and the order you said them in. That includes several things people think of as delivery but which survive as text: a sentence you abandon and restart, a point you make three times, an answer that trails off into "so, yeah, basically" instead of landing on a result. The panel's follow-up, if there was one, and your reply to it are in there too.
What the transcript drops is everything that lives only in sound. Whether you talk too fast, whether your voice rises uncertainly at the end of sentences, how long you paused before starting, how loud you were, whether you smiled: none of it reaches the rating. There is no camera, so no body language, eye contact or clothing either.
Filler words are a grey zone. Whether an "um" appears in the transcript or gets smoothed away is decided by the automatic transcription, not by us, and it is not consistent from one answer to the next. So do not rely on the score to count your filler words. It will not do that reliably.
Where delivery leaks into the score anyway
The structure mark is where delivery counts indirectly. An answer that opens with generalities, jumps to an example, drifts back and stops without a conclusion reads exactly as disorganised on the page as it sounded. The model cannot hear nerves, but it can see the missing thread.
The time limit is the second channel. Each question allows three minutes at most, with a visible countdown, and at zero the answer submits itself. If you wander, you get cut off mid-sentence, and the text is missing the part that mattered, usually the result.
The third channel is intelligibility. Speaking very fast, very quietly or against background noise risks words being transcribed wrongly: one figure becomes another, a technical term turns into an everyday word. The rating reads what was transcribed and has no way of knowing what you meant. Your accent is not graded as such; it matters only if it distorts the transcript, and the report lets you read what actually came through.
The fourth is the follow-up. If the panel asks one, your first answer was evidently too vague. The follow-up is not a penalty in itself, but it sits in the text and shows the information it asked for was missing the first time.
A worked example: same facts, different order
The question comes from the team lead at a mid-sized logistics company: "Tell me about a mistake you made." First version, as spoken and transcribed: "Right, a mistake, that is hard. In the warehouse we had, well it was not only my fault really, we had a wrong stock count, no, first the order was wrong, and then we sort of fixed it, and since then I am more careful."
A plausible rating: 4 out of 10. The reasoning would run along these lines: there is an example, but no clear sequence, the responsibility is hedged, and there is no concrete result. A strong answer would have said what exactly went wrong, what you changed, and how you know the change worked.
Second version, with the same facts put in order and finished: "I entered a reorder in the wrong unit, cartons instead of pallets. It surfaced at the stock count. I sorted out the difference with the supplier and added a check line to the order template. In the months since, that error has not come back." A 7 or 8 out of 10 is plausible here.
The difference is not tone of voice. It is a thread, owned responsibility and a closing sentence, all of which are visible as text. Both versions are constructed to show how the rubric works, not a guaranteed score.
Two cases where the score can mislead you
Case one: you sound calm, assured and pleasant but say little that is concrete. In a real room, presence carries you some of the way. The score only sees thin content and comes out low. Here the score is right, because a Swiss panel will ask for evidence by the second round at the latest.
Case two: your answers are strong and well organised, but you are very nervous, rush your words and look at the table. The score is good because the text is good. In a real interview the same content can come across as weaker. Here the score misses something, and that is what it means to say the score does not know the room.
So if your practice marks are consistently good and real interviews still end in rejections, look at delivery next, not at the content again.
Interviewing in a language that is not your first
Many people practising in English are not native speakers, and many English-language interviews in Switzerland are held by panels who are not native speakers either. There is no separate language mark. The three criteria are substance, structure and fit. Grammar mistakes do appear in the transcript, though, and a sentence that becomes unclear because of its grammar is unclear in content as well. A foreign accent is not scored.
Sessions run in the language you choose: English, German, French or Italian. If the real interview will be in German or French, practise at least a few answers in that language, because phrasing does not carry over word for word and the German voice speaks standard business German, not Swiss dialect. We have not measured how reliably Swiss German dialect is transcribed.
If you are applying from abroad, the questions themselves still follow Swiss conventions, such as salary expectations, notice periods and references, so the content practice is useful whatever your accent.
How to practise delivery separately
Start with the report itself as a mirror. Read your transcript slowly. Abandoned sentences, the same point made three times and a missing ending are right there in black and white, even when no number names them.
Then record the same answer on your phone, on video, without reading from notes, and watch it back with sound. Check three things: pace, how your sentences end, and the first five seconds. It costs nothing and shows more about your presence than any score.
When a lot is at stake, a person is the better mirror. An honest colleague who has sat on interview panels will know at once whether you sound convincing. A coach is worth it for a senior role or after several rejections at the final stage. Many Swiss cities also have Toastmasters clubs, several of them English-speaking, where people practise speaking to an audience.
The sensible division of labour: use the practice interview to make the content strong and matched to the role until your marks are stable, then work on delivery with a recording or a person, using answers that already hold up.
What the score does not promise
The score compares your answer against a rubric: substance, structure and fit for this job, this company size and this role. It does not measure whether you will get the offer. A 9 does not guarantee one, and a 5 does not mean you have no chance.
Real interviews turn on things no transcript shows: chemistry with the team, internal candidates, budget and timing. Treat the score as a work order for your next answer, not as a forecast.
How the conversation and the rating work (live voice, rating from the transcript, 0 to 10 per answer on substance, structure and fit, pass estimate from all transcripts, three minutes per question at most, one follow-up) checked on 16 September 2026 in the implementation of the SwissJobs.app practice interview. The example answers are constructed to illustrate the rubric; the marks given are plausible values, not guaranteed results.
Related questions
What our job index says about the Swiss market
Computed live from our own index, not quoted from a study. Shares only, as of today.
Language the advert is written in
- Deutsch
- 60%
- English
- 23%
- Français
- 13%
- Italiano
- 3%
Of adverts that state a language requirement, the share asking for
- Deutsch
- 70%
- English
- 43%
- Français
- 21%
- Italiano
- 3%
19% posted in the last 7 days · Largest markets: Zürich 18% · Bern 10% · Genève 5% · Basel 5%