Reuters announced this month a new live transcription feature on its Reuters Connect platform: AI-generated, time-coded transcripts of live events as they happen, with real-time clipping that takes a client from a live moment to a published quote in minutes.
Cool, right?
It caught my attention because I've been running a similar experiment in the same direction from my post at the UN.
The candidates to be the next UN Secretary-General have so far sat (or stood at a podium!) through nearly 20 hours of public forums at UNHQ in New York: six three-hour interactive dialogues and a 90-minute televised town hall. (A seventh candidate, the race's newest entrant, faces his dialogue on 19 August.)
I indexed them through AI, and built a site that maps where each candidate stands, issue by issue, in their own words, with every quote linked to the session where they said it.
A project like this used to take a team and a hefty budget. (Think: journalists up all night in the office with empty pizza boxes and lots of coffee.)
I built it alone, with no budget, as a side project on my own time outside my regular reporting schedule.
Building it taught me a lot about verification, and about how transcript tools and AI indexing actually work together. And it turned out to be a good example of knowing which tools to use for what project. It's not one size fits all.
I had two big batches of events to cover — each with different challenges and lots of transcripts.
First: All six candidates each had a 3-hour public hearing (interactive dialogue in UN-speak). Each gave opening remarks, and they each had to field questions for 3 hours from UN Member States. All said and done, six candidates times 3 hours equals 18 hours of transcripts. Most of it was in English, but there was also a lot of French, and strong accents to boot (this is the UN after all!). I normally transcribe short meetings with Otter, but I had doubts Otter would be up to the task for this one. So as I sat through all the public hearings over several days, I took notes — pen and paper, for my own reference. The first batch of transcripts for each candidate I pulled from UNTV. They were "clunky" XLSX files: paragraph numbers and text, navigate by hand. Not an ideal format, but after each 3-hour session I uploaded it to AI and the indexing began.
A day later, I'd pull the same session from the UN's transcript portal, which serves the same automated transcription in a richer format — speaker labels, time stamps, deep links — and upload that too. It made the indexing sharper and attribution easier. But make no mistake: same machine, same guesses. Two copies of one transcript can't check each other. That lesson comes later.
Because here's the deal: machine transcripts are good enough to be very good — and dangerous. Early in the build, I verified eight machine-transcribed quotes against the recordings by ear. Seven contained errors. Three reversed the meaning of what was said. In one, a candidate's "illegal" came out as "legal" — in a quote about humanitarian obligations. In another, a candidate's conditional — that the UN cannot mediate a conflict unless it is engaged and at the table — was flattened into a flat declaration that the UN cannot mediate at all. Same words, missing clause, opposite meaning.
No software flags errors like these, because nothing about them looks wrong. They read cleanly — they just aren't what was said.
So now you might be thinking: if the transcripts produced errors, it shows AI failed, right? Not really.
In the early stages of the build I baked in a rule that cannot be modified by an LLM, no matter how good: no quote is marked verified until a human being, me, has listened to the recording. Until then it carries a visible label telling the reader it came from an automated transcript. AI indexed those 18 hours. The label tells you which sentences a journalist has actually heard. And I told the AI: any transcript quote that seems off — flag it to me. It did.
I said I had two big batches of transcripts. The second one was the Town Hall. This was different, and called for a modified workflow. Six candidates, all on stage at one time, two moderators firing off questions, plus an audience participation segment. All told, 90 minutes. Since it was only 90 minutes, and I was also doing a story for television (my real job), this time I decided to use Otter, because speed was of the essence. I uploaded 20-minute chunks of the Town Hall and fed them directly to AI as it was happening — allowing AI to do real-time analysis of the quotes against what I already had on my site from the public hearings.
Click here to watch the story I did on tight deadline about the Town Hall: https://www.youtube.com/watch?v=lhgE7niQx8k
As expected, the Otter got garbled in several spots. But it was better than nothing.
After the Town Hall was over, AI flagged 29 new quotes from candidates ready for the site. But Otter is far from perfect and should never be trusted with verified quotes. So the AI followed my rules and would not let me put them on the site until all 29 got my ear. That meant start, stop, rewind, verify, and stamp against UNTV. Working alone, that is time I just don't have.
So I went back to the UN transcript portal. It was more accurate than my Otter.
And this time, the second transcript changed everything — because it wasn't a second copy. That gave me two independent transcripts of the same 90 minutes: mine, from a live transcription tool, and the UN's — two different engines. Remember: two copies of one transcript can't check each other. Two different machines can. I had AI cross-reference every quote against both.
Where two machines that had never met agreed word for word, the odds that both made the identical mistake were low. Where they disagreed, a human ear was needed.
The list of quotes needing human verification went from 29 to five. Three hours of work became about 20 minutes. And the editorial standards held.
I verified all five. It mattered. One transcript identified a questioner as Sweden; it was Sierra Leone — which chairs the African Union's committee on Security Council reform, so the mistake would have changed the meaning of the exchange.
Another had a candidate saying "the United States" in an answer where he had simply misspoken; both machines faithfully transcribed the slip, and only listening to the room settled what he meant and how to handle it honestly.
That is the division of labor I now trust. The machines are superb at the work that used to consume the hours: searching, matching, cross-referencing. What was actually said in the room still has to be settled by a person. With transcripts — preferably two of them, from different machines — the job goes from a team and many hours to one person and 20 minutes.
Reuters, to its credit, makes a version of the same point in a separate piece about its AI tools — speed without accuracy is the opposite of what a news organization needs. The open question is where the accuracy comes from. On my site, it comes from a reporter with headphones, checking the small number of discrepancies the machines can't resolve on their own.
The gap between what a global agency can build and what one journalist with the right tools can build has never been smaller. The tools are the same. The judgment still has to show up in person.
The site is here: https://unsgcandidates.orosei.ai/