tl;dr of Italian Transcription Tools And Accuracy
We blind tested to see which was the most accurate Italian meeting transcription tool.
We put tl;dv against three other tools, Google Gemini, Fireflies, and Fathom, to see what would be the best AI notetaker for Italian, and tl;dv came out top with 181 points out of a total of 200.
The areas we scored in were:
- Italian Meeting Transcription And Accuracy
- Real-World Meeting Quality
- Capabilities And Features
- Trust, Security And Value
- Italian Meeting Accuracy Test: Methodology
We used the same Italian recordings and scored every transcript and summary using both AI and a native Italian speaker for our assessment. tl;dv came out by a wide margin, followed by Fireflies, then Gemini, with Fathom in last place.
Italian meeting transcription accuracy can vary a lot, despite what the marketing of different meeting tools suggests. Based on our recent testing of different language transcriptions from tools, we often find that what a tool says it can do and the actual output can be very different.
So we tested tl;dv’s Italian transcription capabilities against three other tools that offer Italian as a marketed option: Google Gemini, Fathom, and Fireflies.
The results were interesting, with some of the tools struggling to capture some of the basics when it came to naturally spoken, native Italian, and through all our testing criteria tl;dv consistently came out on top.
Scoring was done independently by two LLMs and confirmed by a native Italian speaker, none of whom knew which tool produced which output.
Below is the overall scoring for each section, and we break down each area of testing further to explain how we concluded that the best Italian meeting tool is tl;dv.
| Tier | Max | tl;dv | Google Gemini | Fireflies | Fathom |
|---|---|---|---|---|---|
| Transcription & accuracy | 65 | 61/65 | 34/65 | 34/65 | 31/65 |
| Real-world meeting quality | 45 | 41/45 | 32/45 | 32/45 | 21/45 |
| Capabilities and features | 72 | 61/72 | 37/72 | 58/72 | 50/72 |
| Trust, security and value | 18 | 18/18 | 15/18 | 18/18 | 12/18 |
| Overall score | 200 | 181/200 | 118/200 | 142/200 | 114/200 |
| Rank | 1 | 3 | 2 | 4 |
Italian Meeting Transcription & Accuracy
These were scored by LLMs (Anthropic’s Claude and OpenAI’s ChatGPT) and then confirmed based on a blind native-speaker assessment, with the tool names hidden.
In our first round of testing, we looked at the actual quality of the transcription itself across each of the AI meeting notetakers: Can they write down exactly what was said in Italian?
Each tool was given the same set of Italian speech, and the output was marked on specific measures from general accuracy at word-level, to elements such as proper noun recognition, how it managed numbers, and specific technical jargon.
In this part of the test, tl;dv came out on top with 61 out of 65 points. The other three tools finished quite closely to each other, in the mid-thirties. This was quite a gap, and below you can see the breakdown of each element we tested and the scoring based on two blind rounds of LLM testing and a native-speaker blind test.
| Metric | How scored | tl;dv | Google Gemini | Fireflies | Fathom |
|---|---|---|---|---|---|
| Language accuracy | Blind native-speaker severity rating on in-language accuracy | 18/20 | 10/20 | 9/20 | 8/20 |
| Language-specific handling | Diacritics, punctuation, regional variants, code-switching | 18/20 | 12/20 | 13/20 | 12/20 |
| Word error rate scoring | Computed against an official transcript or reference text | 5/5 | 3/5 | 2/5 | 1/5 |
| Entity detection | Names, companies and places across the cast | 5/5 | 3/5 | 2/5 | 2/5 |
| Numbers, dates and currency | Figures, dates and amounts formatted correctly in-language | 5/5 | 2/5 | 2/5 | 4/5 |
| Technical term raw recognition | Industry terms and acronyms before custom training | 5/5 | 1/5 | 2/5 | 2/5 |
| Punctuation and segmentation | Sentence breaks and paragraphing in test-run output | 5/5 | 3/5 | 4/5 | 2/5 |
| Transcription & accuracy subtotal | 61/65 | 34/65 | 34/65 | 31/65 |
How Accurate Is Each Tool On Everyday Italian?
Each tool did a fairly good job at capturing general day-to-day Italian. None produced a transcript that was hard to follow or littered with large errors, which isn’t the case for some of the other languages we have tested, including French and German.
tl;dv was the strongest by a fair margin. There was a specific section of audio, centered around finance, which we were able to check directly against a confirmed, published reference text. tl;dv was able to track the original format almost word-for-word, with some discrepancies, as the published text had been cleaned up, whereas tl;dv was able to keep it verbatim.
Gemini and Fireflies were also able to keep up fairly well on general sentences. There were some very specific errors such as “aduzione” for “adozione“, “traduzionali” for “tradizionali“, or “veramente” where the speaker said “meramente“. Fireflies also turned the phrase “capitale di rischio” into “capirebbe il rischio” which is not a real term.
Of the four tools tested, Fathom was the one that produced the most errors. In some areas, the AI appeared to try to guess specific phrases, turning “Basti pensare” into “Passi a pensare” and “manifattura” into “manufattura“. There was also a sentence about patient finance and risk capital that Fathom didn’t capture at all.
Fathom was the weakest when it came to raw Italian accuracy.
Which Tools Get Italian Acronyms And Technical Terms Right?
Meetings are regularly filled with incredibly technical terms and acronyms, and are likely to be incredibly important to the overall context of the meeting itself. In our testing on Italian transcription, each clip was selected based on the amount of technical jargon to see how easily it was picked up by the transcription tools. One in particular, centered on healthcare, had many terms that were Italian shorthand for specific things. Elements included EHDS regulation, the NIS 2 directive, and also factored in a range of terms on cybersecurity and interoperability.
Out of all four tools, only tl;dv was able to keep them correct, capturing EHDS and NIS 2 correctly each time, without loading any custom vocabulary to the interface.
For comparison, Gemini turned “NIS 2” into “IS2” and rendered “EHDS” three different ways in the same passage, as “HDS“, then “HS“, then “HD“.
Similarly, Fireflies completely omitted “NIS 2″ and changed “EHDS” to “HDS“.
Fathom completely missed them, changing “NIS 2” to “URIS 2” and “EHDS” to “HTS“.
These can be rectified, and many of the tools offer custom vocabulary, but capturing them correctly from fluent Italian in the first place matters more.
On these particular lines, tl;dv scored full marks.
How Well Do They Handle Numbers And Dates In Italian?
Dates and figures are where a single error flips the meaning of a whole sentence, so we tracked them on their own. The healthcare speech turned on one number: the 2029 deadline for the European Health Data Space. tl;dv and Fathom both caught it cleanly. Gemini truncated it to “202”, which reads as a typo or a different figure entirely, and it did so more than once. Fireflies spelled it out and got it wrong, writing “duemilaventidove” rather than “duemilaventinove“.
tl;dv scored five, Fathom four, and Gemini and Fireflies two apiece.
Did The AI Notetakers Get Italian Names Right?
In one of the clips, there was a large roll-call of names which allowed us to identify how well the tools were able to manage Italian names. This is really key in areas such as sales, when dealing with prospect names and key stakeholders with a connected CRM.
The list of names from this included Gratteri, Frattasi and Melillo to Valensise, Morrone, Caravelli, Di Monte and Costanzo. The tools differed greatly between the outputs.
tl;dv was able to manage all the list correctly and was the only tool able to accurately capture “Angelo Costanzo” correctly rather than “Costanto“.
Fathom cleanly captured “Giovanni Caravelli” where two of the other tools added a stray “Gianni“, but it turned “Simona Di Monte” into “Simone Di Monti“.
Gemini and Fireflies both missed many of these and then captured “Battarella” rather than “Mattarella“, an error Fathom also repeated.
tl;dv took five points here; the rest ended up on two or three after accounting for the errors.
Real-World Meeting Quality
Transcription accuracy is at the core of the quality of the output, but how that output transfers into other areas of daily working life is also something worth considering. The output of all four of the tools tested also comes with elements such as summaries, and these need to be good enough to send out to others for any AI meeting assistant to be worth using.
In this section of our testing of Italian meeting transcription, we looked at elements such as hallucination rates, summary quality and action points.
We also looked at diarization quality, which held up on the tools we could test, with Gemini needing to be prompted before it separated the speakers at all. Fathom could not be scored here, as it only live-records and our live runs were single-source.In this section, tl;dv scored the highest marks.
| Metric | How scored | tl;dv | Google Gemini | Fireflies | Fathom |
|---|---|---|---|---|---|
| Diarization quality | Correct speaker count and turn attribution vs known cast | 9/10 | 7/10 | 8/10 | 0/10 |
| Behavioral stability | Behavioral stability across session types | 9/10 | 6/10 | 5/10 | 4/10 |
| Summary quality | Usefulness of the summary and whether it stayed in the source language, with allowances for loanwords | 5/5 | 4/5 | 3/5 | 5/5 |
| Hallucination / insertion rate | Invented, looped or duplicated text not present in the audio, with minor deductions where mishearings altered meaning | 10/10 | 9/10 | 9/10 | 4/10 |
| Action item extraction | Quality of tasks and follow-ups pulled from the meeting | 4/5 | 2/5 | 4/5 | 3/5 |
| Auto chapters / sectioning | Does the summary break the meeting into useful sections | 4/5 | 4/5 | 3/5 | 5/5 |
| Real-world meeting quality subtotal | 41/45 | 32/45 | 32/45 | 21/45 |
Which Tools Hallucinate Or Invent Text?
Hallucination is one of the biggest concerns that people face when dealing with anything to do with AI. If a tool invents and fabricates wording, actions, or anything else, it can lead to serious errors and potentially bigger issues. A single misheard word can be irritating, but it can normally be easily spotted; whole sections of invented text can be a lot worse.
In our review, tl;dv invented nothing at all across all of the recording sessions. Each element was true to the baseline and didn’t create anything different from what was said.
Gemini and Fireflies also managed to avoid inventing and inserting anything different or made-up when it came to transcribing in Italian. Both scored just short of full marks, docked for mishearings that shifted meaning rather than for anything invented.
Fathom, however, inserted whole phrases that never appeared in the audio. “Grazie per aver guardato” appeared in one of the transcripts, which is “thanks for watching”. Other sentences that popped up were “Grazie a tutti” and a stray “Siamo capaci di fare noi“. Also mentioned in our first section, it dropped a full sentence in the finance run.
This is concerning as a fabricated line in an otherwise good transcript is potentially the worst failure that any kind of AI notetaker can have. It’s fairly unlikely that a person will proofread a summary before it is sent out if it reads well.
Which Tool Writes The Best Italian Meeting Summary?
This is the area where some of the tools that have struggled did quite well. Despite having the worst output on transcription, Fathom’s summary was one of the best. It was organized well into purpose and key topics, and errors from the transcript, such as EHDS, appeared correctly in the summary, which is strange when they were so wrong in the raw capture.
tl;dv was able to match it for usefulness and correctness. Gemini also produced a solid summary, but not quite as strong as the others. Fireflies wrote incredibly detailed summary notes; however, they ran incredibly long and were hard to skim-read.
The summary sits above the raw transcript, so it’s likely that the AI can “tidy it up” at this stage. But it does flag that solid summaries may not mean that the actual data that it recorded on the raw transcription actually matches up.
How Consistent Is Each Tool Across Meeting Types?
One of the areas that we looked at was behavioral stability. The summary versus transcription when it came to Italian showed that there can be some differences in how effectively the tools can manage things over different meetings.
tl;dv was able to keep solid stability across all the testing, carrying a 9 out of 10 score. Gemini was also fairly consistent, if a little middling with the details. Fireflies had a few wobbly bits, struggling on the healthcare acronym terms. Fathom was the worst performer, doing well when the topic was justice, struggling with finance terminology, and then inventing text on the healthcare subject.
Capabilities And Features
Aside from the Italian transcription and summaries that each tool generates, each of the tools differs a lot in terms of how it packages its overall ability. tl;dv for example is more than an AI meeting assistant, but is a full suite that offers organizational memory, coaching, layer-upon-layer of value, and can integrate directly with CRM systems, Claude and ChatGPT.
In this section we take a look at how each tool scored when it comes to the various opportunities to extend and use the data, features and the overall quality of some of the elements such as processing speed and timestamp accuracy.
In this area tl;dv led again with 61 out of 72, followed by Fireflies with 58, Fathom and Gemini were below this.
| Metric | How scored | tl;dv | Google Gemini | Fireflies | Fathom |
|---|---|---|---|---|---|
| Speaker naming out of the box | Auto-names real speakers on the platforms it supports | 5/5 | 5/5 | 5/5 | 5/5 |
| Voice printing | Availability of voice-print training for the user’s own voice | 5/5 | 0/5 | 0/5 | 0/5 |
| Bot-free recording | Records via system audio without sending a bot into the call | 5/5 | 5/5 | 5/5 | 5/5 |
| CRM sync | Native and auto-sync | 3/3 | 0/3 | 3/3 | 3/3 |
| Custom notes / templates | Customizable summary formats vs a fixed output | 3/3 | 0/3 | 3/3 | 3/3 |
| Custom vocab / entity training | Teach industry terms and acronyms | 5/5 | 0/5 | 5/5 | 5/5 |
| Italian UI localization | Whether the product interface itself is available in Italian | 0/5 | 5/5 | 0/5 | 0/5 |
| Integrations breadth | Slack, calendar, Zapier, API | 3/3 | 3/3 | 3/3 | 3/3 |
| Processing speed | Time from meeting-end to finished transcript | 3/3 | 1/3 | 3/3 | 3/3 |
| Filler-word tracking | Tracks ehm, cioè, allora without stutter-doubling. Allows for full visibility of spoken transcripts rather than over-smoothing | 3/3 | 0/3 | 0/3 | 0/3 |
| Timestamp accuracy | Spot-check that timestamps land on the right moment | 3/3 | 3/3 | 2/3 | 0/3 |
| Translation availability | Can it translate the meeting notes, and into how many languages | 3/3 | 0/3 | 3/3 | 3/3 |
| Search within transcript | Search across a meeting and across the library | 3/3 | 3/3 | 3/3 | 3/3 |
| Transcript editing UI | Can you correct the transcript easily after the fact | 3/3 | 3/3 | 3/3 | 3/3 |
| Export formats | SRT, VTT, TXT, DOCX and similar | 0/3 | 3/3 | 3/3 | 0/3 |
| Live / real-time transcript | Is a transcript shown live during the meeting | 0/3 | 3/3 | 3/3 | 3/3 |
| Meeting platform coverage | Zoom, Meet, Teams, Webex coverage | 3/3 | 0/3 | 3/3 | 3/3 |
| Mobile app capture | Can it record in-person meetings via a mobile app | 3/3 | 3/3 | 3/3 | 0/3 |
| Native MCP server | Native first-party server letting AI assistants query the meeting library | 5/5 | 0/5 | 5/5 | 5/5 |
| Speaker label editing | Can you rename and reassign speakers after the fact | 3/3 | 0/3 | 3/3 | 3/3 |
| Capabilities and features subtotal | 61/72 | 37/72 | 58/72 | 50/72 |
Recording And Speaker Handling
All four tools have the ability to auto-name speakers on the main meeting platforms, so the set is pretty even. The area where some tools pull away, notably tl;dv, is how deep they are able to do this.
Of the four tools tested, only tl;dv has the ability to offer voice-printing training, where it can learn your voice once and recognize you in future meetings if you opt in for this.
Three of the tools allow for speaker relabelling in the interface, tl;dv, Fireflies and Fathom. Gemini, because of the way that it processes and displays transcripts (in Google Doc format), this technically can be done but this is a manual process. As a result, across the board, Gemini was the most boxed-in of the four on speaker controls, losing points on label editing and reassignment.
If you are in a team that shares meeting recordings across a business, they need to have clean and correctly attributed speakers, so this is a strong point of separation, and one where Google’s tool falls behind.
Customization And Vocabulary Training
This ties back to the acronym and naming failures we found in the Italian transcription in the first section of scoring. Any issues could be rectified by plugging the industry-specific terms, difficult names, and acronyms into the custom vocabulary setting of the tools.
tl;dv and Fireflies let you teach the tool your own vocabulary on the plans we ran, Fathom only on its top Team plan, and Gemini does not at all. So as a result what Google’s Gemini hears is what you get, and there is no way to fix or influence what the tool hears. tl;dv also supports custom summary templates rather than a single fixed format, so the output can match how your team already writes its notes.
Filler Word Tracking
One feature that protects the integrity of a true transcript is filler-word visibility. Rather than cleaning them out and smoothing the record, tl;dv is the only tool that keeps filler words in the transcript itself, so you can see what was actually said. Fireflies counts fillers too, but as a separate speaking-analytics metric rather than keeping them in the transcript.
Trust, Security And Value
In our scoring, this final section is based around one of the most important elements when it comes to recording meetings, namely, what happens to your recordings once a meeting ends, and what each tool costs to run?
This carries a lot of weight in Italian because a meeting recorded in Italy sits under EU rules such as GDPR, which dictates where your data is hosted, retention attributes, and whether or not it is used to train someone else’s AI model.
Each tool is open and transparent about their pricing, and do not require a sales call. Each tool also carries a range of security certifications and accreditations that you would expect with a business tool. Generally there are good scores for each tool we looked at.
Tools mainly begin to lose points on areas like being able to select where your data is held. EU regional hosting, for example, is something that Fathom, an American company, does not offer, although Fireflies, also US-based, does offer it on its top plan.
Gemini loses points because there is no free option. Even with a paid-for Google workspace account, the ability to use Gemini for notetaking in Italian, or any of the other languages it supports, is ringfenced to certain paid-for packages.
The biggest flag is about AI training. Fathom has, deep in its settings, a toggle to switch this off. You need to find this and switch this off, because the tool will automatically use de-identified meeting data to improve its own AI. For a single user that may not matter much, but it is exactly the kind of default a business would want to know about before a team-wide rollout, especially one handling sensitive or regulated conversations.
| Metric | How scored | tl;dv | Google Gemini | Fireflies | Fathom |
|---|---|---|---|---|---|
| Data residency / regional hosting | Regional hosting options, e.g. EU hosting on demand | 3/3 | 3/3 | 3/3 | 0/3 |
| Security and compliance | SOC 2 Type II and GDPR compliance | 3/3 | 3/3 | 3/3 | 3/3 |
| AI training on user audio | Does it avoid training AI on your audio (no training scores full marks) | 3/3 | 3/3 | 3/3 | 0/3 |
| Data retention controls | Control over how long recordings and transcripts are kept | 3/3 | 3/3 | 3/3 | 3/3 |
| Price transparency | Plan prices are published rather than quote-only | 3/3 | 3/3 | 3/3 | 3/3 |
| Free tier / limits | Free plan availability (a free trial alone scores 0) | 3/3 | 0/3 | 3/3 | 3/3 |
| Trust, security and value subtotal | 18/18 | 15/18 | 18/18 | 12/18 |
Italian Meeting Accuracy Test: Methodology
Our comparison is built on a controlled, like-for-like test designed to give every tool the same conditions.
The Test Set
For testing we selected three Italian recordings from the public archive of Radio Radicale. Each one was selected because it offered detailed information and had areas that would push the tools in various weaknesses around terminology, speed of speakers, and other areas. The first was the opening of Banca d’Italia and European Investment Bank conference on finance and artificial intelligence, delivered by the Governor, Fabio Panetta. This one in particular had detailed economic language in Italian, and was matched with a published official text for reference.
The second was a summit on health data and artificial intelligence, from members of the Senate, chosen for a lot of acronyms and medical terminology mixed with technological jargon.
The third clip selected was the opening of the Court of Appeal of Naples conference on AI, justice and security, chosen for legal terminology, technological terminology, and a good selection of names to test the proper noun functionality.
Every clip was condensed to under ten minutes, and each tool received the same audio for fairness.
The Review
For blind testing, each transcription and summary was stripped of its tool name, links and any other identifiable data and relabelled with a unique indicator. This meant that neither the LLMs or the native-speaker would be able to identify which transcript or summary belonged to which tool. These anonymized outputs were then sent to two LLMs, Anthropic’s Claude and OpenAI’s ChatGPT, which ranked them independently on the first two sections of the testing.
We then enlisted our native Italian speaker Pietro Giucastro, who checked the trickiest phrasings and where the tools gave different outputs, by ear, with the original videos and the anonymized outputs to confirm the original machine scoring. All three assessments landed on the exact same order, confirming that the ranking holds up across three different tests, on three different judgments across the board.
The Tool Set
We selected tl;dv along with three other tools to score against; each one was selected as they offer Italian transcription clearly in their marketing and would be tools that Italian teams would be selecting from.
- tl;dv: EU-based, internationally run tool and one of the most widely used notetakers for multi-language meetings. We ran it on the Business plan, but it uses the same transcription engine underneath as the free tier.
- Google Gemini: included because it comes bundled with Google Workspace, which makes it the notetaker many Italian teams already have switched on by default.
- Fireflies: US-based and one of the most popular standalone AI notetakers, and a common pick among Italian users.
- Fathom: US-based, which a lot of smaller teams start out on, so it represents the free end of the market.
When we tested, we ran the tool on either a free or paid account, so the results will reflect what an ordinary user would get. The live runs captured only a single audio source, so where a tool could take an uploaded file, we used a separate multi-speaker run to score diarization quality.
Engine & Plan Breakdown
While all tools offer Italian transcription, they do not run on the same engine, which explains why the results differ so much. tl;dv’s testing used a dedicated speech engine from ElevenLabs. Gemini uses Google’s own model. Fireflies runs on OpenAI’s Whisper, based on current documentation, with an additional layer on top. Fathom licenses a third-party engine that is not named publicly. Below you’ll see the engines and what plan we tested on.
| Tool | Underlying engine/vendor | In-house or licensed | Engine type | Plan |
|---|---|---|---|---|
| tl;dv | ElevenLabs | Licensed | Dedicated ASR | Business |
| Google Gemini | Google Gemini | In-house (Google) | LLM | Standalone Gemini app (Business Starter account) |
| Fireflies | OpenAI Whisper (plus in-house LLM layer) | Licensed | Dedicated ASR | Pro |
| Fathom | Third-party, not publicly specified (AI sub-processors: OpenAI, Google, Anthropic) | Licensed | Not disclosed | Free |
Scope & Caveats
With any testing, we will always try to ensure that it’s as balanced and scientific as possible; however, there can be limits.
- All three test clips were single-source audio for the live test runs; we ran a separate multi-speaker pass on the tools that accept uploaded files to score diarization quality.
- The reference material we used is published directly from the publisher, but is not a true verbatim account; it is cleaned, so there was an element of flexibility built into the scoring based around this.
- Fathom also only live-records and does not offer the opportunity, when we tested, to upload directly to the platform, so it could not be scored on diarization and takes a zero on that measure.
- All the capabilities, trust, and security were based on current public product documentation at the time of testing.
What Is The Best Meeting Transcription Software For Italian
For the best Italian transcription accuracy tl;dv was the best performer, capturing the Italian meetings in greater detail and accuracy than the other tools tested. It won each tier that measured the quality and was 39 points ahead overall. Where it won, it was able to handle easy sentences but was also able to handle the trickier parts that the other tools failed on. This was in areas such as names, proper nouns, and general entity detection, and it did not hallucinate or invent any text. This means that the output can be trusted to be the truest account of what has been said in Italian compared with the other tools we scored.
It also won across all three judges, two LLMs and one native speaker, all testing blind.
If you hold meetings in Italian and care about how to record, and more importantly use, the output of your meetings without the need to check, tl;dv is the clear winner here.
With solid features, customization, and a native MCP you can plug directly into Claude or ChatGPT, it is the ideal tool for Italian meeting transcription that will take your meetings and turn them into actionable, useful output for you to use.
Try tl;dv for yourself today on your next Italian meeting and see for yourself.
FAQs About Italian Meeting Transcription in 2026
Which AI notetaker is most accurate for Italian meetings?
tl;dv scored 181 out of 200 across three blind-scored Italian recordings, well clear of second-placed Fireflies. Gemini came third, Fathom last. On raw transcription accuracy alone tl;dv scored 61 out of 65 while the other three finished in the mid-thirties. Two LLMs and a native Italian speaker, all scoring blind, returned the same order.
Do AI notetakers transcribe Italian acronyms correctly?
Only tl;dv did. Given healthcare audio containing EHDS and the NIS 2 directive, it captured both correctly every time with no custom vocabulary loaded. Gemini rendered EHDS three different ways in one passage: HDS, HS, HD. Fireflies dropped NIS 2 entirely. Fathom turned it into “URIS 2”.
Do AI notetakers get Italian names right?
Does AI transcription hallucinate in Italian?
Fathom did. It inserted phrases absent from the audio, including “Grazie per aver guardato”, a YouTube-style sign-off with no source in the recording, plus “Grazie a tutti” and a stray “Siamo capaci di fare noi”. It also dropped a full sentence. tl;dv, Gemini and Fireflies fabricated nothing.
How was this Italian transcription test scored?
Blind. Every transcript and summary was stripped of tool names, then scored independently by two LLMs, Claude and ChatGPT, and confirmed by native Italian speaker Pietro Giucastro. All three returned the same ranking. Audio came from the public Radio Radicale archive. tl;dv ran the test and published the full method.



