AI Notetakers in 2026: What Changed, and What to Check Before One Joins Your Next Call

AI Notetakers in 2026: What Changed, and What to Check Before One Joins Your Next Call

On September 14th, 2026, Superhuman acquired Fathom. That was the fourth structural move in the AI notetaker market since December 2025, and the third since the start of August 2026. If you are choosing an AI notetaker for your team, the useful question is no longer which one writes the best summary. It is which one will still exist in a year, and what it does with the recording in the meantime.

We use these tools ourselves, and we build products on the same frontier model APIs that sit underneath them. Our own calls are the kind these tools handle the worst. Being in the middle of it is why three things changed how we evaluate the category in 2026, and none show up in a feature table.

Nine months that reset the category

Date What happened
Dec 5, 2025 Meta acquires Limitless, the AI pendant company, announced by founder Dan Siroker and reported by TechCrunch. Hardware sales end.
Mar 25, 2026 Granola raises $125 million at a $1.5 billion valuation, up from $250 million at its previous round.
Jul 30, 2026 Granola is sued in a proposed class action for recording someone who never used it.
Aug 5, 2026 Wispr Flow, a dictation company, ships a meeting notetaker that does not join the call. Windows followed on Sep 15.
Aug 13, 2026 A federal judge lets the core privacy claims against Otter.ai survive a motion to dismiss.
Aug 19, 2026 Calendly, a scheduling company, ships Calendly Notetaker after more than 100,000 beta recordings.
Sep 14, 2026 Superhuman, an email company, acquires Fathom, which TechCrunch reports had over 400,000 monthly active users.

Read that as one sequence, not seven headlines. What it suggests is that capture is turning from a business into a feature of something else.

Why are scheduling and email companies suddenly building notetakers?

Because recording is no longer the hard part. The value sits in what happens before and after the call, and any company that owns one of those moments can attach capture for almost nothing. Calendly owns the booking. Superhuman owns the inbox. Google, Zoom and Microsoft own the room.

That is a distribution play rather than a product one. Fathom's founder said as much when the deal was announced, that running standalone would have meant building what Superhuman already has.

The independents still standing never played the growth game. Fireflies.ai reached a $1 billion valuation in June 2025 through a tender offer, no primary capital raised since 2021, profitable since 2023.

So what does a buyer do? A standalone notetaker is a bet on an acquisition target. That is not necessarily a bad bet, but it does mean the terms covering your recordings can change hands without you.

Does removing the bot make a notetaker private?

Not on its own, and "botless" covers two different things. A platform like Google Meet or Zoom records natively and still announces it. A third-party tool capturing system audio on one machine announces nothing. Only the second removes the signal other people rely on, and neither keeps the audio local.

This is the most common confusion we see, and it separates into two questions.

The first is how the audio gets reached at all, and there are three answers. A third-party bot dials in and shows up in the participant list. The platform records natively, because Google or Zoom already holds every stream. Or a third party captures system audio on one machine and never touches the meeting. Only the first is visible to everyone else, which is why Microsoft now treats uninvited bots as a security problem, placing detected ones in the lobby "irrespective of the lobby configuration for the meeting" (Microsoft Learn), with an admin control that auto-blocks them due in October.

Where the audio goes after capture matters more. Granola is the clearest case, and it cuts both ways. Nothing joins your call, and its security page names "transcription providers (like Deepgram and Assembly) and AI providers (like OpenAI and Anthropic)" and states that "Granola trains on your anonymized data," with an opt out and training off by default for enterprise. I have not seen a clearer disclosure from a notetaker vendor.

It is also the subject of Chamberlain v. Granola, Inc., a proposed class action filed in July 2026 by someone who never used Granola but was recorded when another participant did. It pleads the federal wiretap statute plus California wiretapping and confidential communications counts, and alleges the chat notification and on-video watermark both default to off. It quotes Granola's own marketing as having said other people in the room would not know it was there.

Honest documentation and a visible consent signal are not the same thing, and only one is what the other people on your call experience. Local capture is also not local processing. So the question to put to a vendor is not whether a bot shows up in the participant list. It is which third parties touch the audio, and what they are allowed to do with it.

Diagram comparing three ways an AI notetaker gets meeting audio. A third-party bot appears in the participant list, native recording shows an indicator, and local system capture shows nothing to other participants.

What did the Otter.ai ruling actually decide?

On August 13th, 2026, Judge Eumi K. Lee granted in part and denied in part Otter's motion to dismiss. Federal wiretap, California Invasion of Privacy Act and Illinois biometric privacy claims all survived. The reasoning did not hinge on the act of recording. It hinged on what the vendor did with the recording afterward.

The operative sentence in the decision is worth reading closely. "Because Plaintiffs plausibly allege that Otter independently collects, retains, and uses communications for its own commercial purposes, they have sufficiently alleged that Otter is a third-party eavesdropper under section 631."

Retention and commercial use moved the vendor from an extension of the user to an independent party to the conversation. Those are architecture decisions, made long before they become legal exposure.

A second thread runs through Illinois biometric law. Cruz v. Fireflies.AI Corp., filed in December 2025, targeted voiceprints taken from people who "never created Fireflies accounts, never agreed to Fireflies' Terms of Service, and never executed any written consent." Cruz was dismissed without prejudice in March 2026, and the theory now sits before the Northern District of Illinois in Fricker v. Fireflies.AI Corp. The alleged violation is speaker identification. If your notetaker labels Speaker 1 and Speaker 2, it is arguably generating biometric identifiers for everyone on the call, including guests who agreed to nothing.

None of this has been decided, and I am not a lawyer. What interests me is less how these cases come out than what they are circling. If the exposure keeps landing on retention, identification and notice rather than on the microphone, then this is an architecture question rather than a policy one, and the three capture models above do not sit in the same place on it.

The accuracy number in the marketing is not your accuracy

Published transcription accuracy is almost always benchmarked on read, American-accented speech. Real meetings are neither.

AfriSpeech-MultiBench evaluated 19 speech recognition systems across more than 100 African English accents and 79 hours of audio. Whisper-large-v3 scores 2.01 word error rate on LibriSpeech and 26.49 on African-accented English, roughly thirteen times worse on the same model. The pattern holds across vendors, with AWS Transcribe at 32.77 and Google Chirp V3 at 35.03.

The detail that should worry anyone relying on a summary is named entities. Most systems in that benchmark exceeded 40 percent error on names, numbers and financial commands, most landed between 60 and 78 percent, and the worst exceeded 84 percent. Names, companies and figures are what a summary exists to capture, and they fail most often.

Aggregate scores hide this too. A September 2026 study of English and Yoruba code switching found two systems with near-identical word error rates performing very differently on the switches themselves. One number cannot tell you whether a tool works for your team.

Then there is invention rather than error. The Careless Whisper study at ACM FAccT in 2024 found that "roughly 1% of audio transcriptions contained entire hallucinated phrases or sentences which did not exist in any form in the underlying audio," and that 38 percent carried explicit harms. They clustered around speakers with longer pauses, and meetings are full of pauses.

We run engineers in West Africa and in the United States on the same calls. This is not an abstract fairness concern for us. It is what a transcript of a normal Tuesday looks like before anyone edits it. If your team has accents these benchmarks do not cover, or names these models were not trained to expect, treat every AI-generated summary as a draft written by someone who was half listening.

Bar chart of word error rate. Whisper large-v3 scores 2.01 on LibriSpeech and 26.49 on African-accented English, while AWS Transcribe scores 32.77 and Google Chirp V3 scores 35.03. Lower is better.

What should you ask a vendor before letting it into a client call?

Six questions, in the order they matter. Most sales pages answer the first two.

  1. Retention. How long do you keep audio and transcripts, as a number of days rather than "as long as necessary"? Otter's privacy policy still uses the second formulation.
  2. Training. Do you train on customer data? Is it on or off by default, can we turn it off ourselves, and does any of that change by plan tier?
  3. Subprocessors. Name every third party that touches the audio or the text, then compare that against the marketing.
  4. Speaker identification. Do you generate voiceprints, and what is the destruction policy for those specifically?
  5. Notice, and its defaults. How are non-account holders told recording is happening, is that notice on by default, and what happens when someone declines?
  6. Termination. If you are acquired, what happens to the archive, and can we export and delete everything first?

Add one that is not about the vendor. Decide, in writing, which meetings may be recorded at all. The most instructive incident here was not a breach. In September 2024 an investor meeting was recorded by Otter, and the transcript was emailed to the outside participant afterward, including "hours of their private conversations" (Entrepreneur). No attacker, no vulnerability, no misconfiguration. The product had no state for "the meeting is over," and default-on distribution did the rest.

If you are signing off on a notetaker across an engineering organization, this is a data handling decision wearing a productivity tool costume. Write down three things first. Where the audio is processed, what is retained and for how long, and how badly the model performs on the accents on your calls. If you want a second pair of eyes on that, contact Ogun Labs.