Is Falcon ASR safe? Where your audio goes in the demo
Falcon ASR, TII’s new speech-to-text model, launched on 7 October 2026. What its demo does with your recording, what is documented and what is not.
Short answer: Falcon ASR is a new speech-to-text model from the Technology Innovation Institute in Abu Dhabi, and today the only way to try it is a public web demo. That demo sends your recording to TII’s own servers to be transcribed. The demo’s documentation is unusually specific about what the page itself keeps. It says nothing about what happens to your audio once it reaches TII’s API, and that is the part that matters.
So it is reasonable to try with a clip you would be happy for a stranger to hear. It is not yet something to feed meeting recordings, customer calls or anything with other people’s voices in it. Here is what is written down, and where the gaps are.
What was launched, and when
TII announced three models together on 7 October 2026, according to the UAE’s state news agency WAM: Falcon-Emirati, a language model for the Emirati dialect, Falcon-OCR-Arabic for reading Arabic documents, and Falcon ASR. TII’s launch post on Hugging Face, dated the same day, describes Falcon ASR as a 1.6 billion parameter model that turns speech into text in Arabic, including Emirati and other Gulf dialects, as well as English, French, Spanish and Portuguese. All five languages share the same weights, and you do not have to tell it which language you are speaking.
The numbers TII publishes are word error rates, the share of words a transcript gets wrong. It reports an average of 20.92% across six public Arabic test sets and 5.74% across seven English ones. Those are TII’s own figures. In plain terms, about one word in five is wrong on hard Arabic audio and about one in twenty on English. That is competitive, according to TII, and it still means every transcript needs reading. Gulf News lists the intended uses as customer service, meeting and interview transcription, subtitles and accessibility tools.
The model builds on TII’s earlier audio research, described in its Falcon3-Audio paper. If the difference between a model and the app wrapped around it is fuzzy, AI agent vs LLM covers it quickly. It matters here, because the app is the part that handles your data.
You cannot download it yet
Many of TII’s earlier Falcon models were released as open weights you could run on your own machine. Falcon ASR, so far, is not. The launch post points to a hosted demo for clips up to 60 seconds and says API access and native apps are planned. When we checked TII’s Hugging Face organisation on 9 October 2026, there was no Falcon ASR model to download.
That changes the privacy question completely. A model running on your own laptop never sends audio anywhere. A hosted model always does. Until TII publishes weights, the second case is the only one there is.
What the demo does with your recording
The demo’s code is public, and its README spells out the flow. It describes the Space as a front end for an externally hosted Falcon ASR deployment and says the model and API remain on TII’s servers. Your upload or microphone recording is checked for length and size, converted to a standard audio format, then sent to that API. The transcript comes back with timings for each word.
- The page itself does not ask you to sign in. Usage limits are enforced per IP address, in memory.
- Its usage analytics, by the README’s own account, contain no Hugging Face account identity, no audio, no transcripts, no IP addresses and no email addresses.
- Uploaded audio sits in the demo’s temporary cache. The README says that cache is checked hourly and files older than 24 hours are removed. Gradio’s documentation explains why such a cache exists at all: files are copied there so one user cannot overwrite another’s.
- A noise gate on microphone recordings is on by default. The README is candid that it does not isolate the speaker, so nearby voices above the threshold may still be transcribed.
That last point is easy to miss. Record in an open office and a colleague’s side conversation can end up in the transcript, and on TII’s servers.
What is not documented
Everything above describes the demo page. None of it describes the API the audio is sent to. Neither the launch post nor the README states:
- how long TII’s servers keep the audio or the transcript,
- whether recordings sent through the demo are used to train or evaluate future models,
- where those servers are, or which country’s data law applies to them,
- who to contact to have a recording deleted.
None of that is evidence of anything bad. A research demo launched this week often has no product privacy notice yet, and the README reads like it was written by people who care about the question. But a gap is a gap. Do not fill it with an assumption in either direction.
How to use it sensibly
- Use clips you would be comfortable posting publicly: a sample you recorded for the purpose, a public speech, a podcast you have rights to.
- Keep other people’s voices out unless they have agreed. A recording of a colleague is their personal data as much as yours.
- Do not upload client calls, interviews or meeting recordings until TII publishes terms that say how long audio is kept and whether it is trained on.
- Read the transcript before you use it. With roughly one Arabic word in five wrong on hard audio, a name or a number can come back changed. How to check an AI answer when you are not the expert works just as well for transcripts.
- If you want to use it for customer service, wait for the API and its terms. Generative AI in customer service covers what to ask a vendor before any tool touches customer conversations.
Voice tools are arriving from every direction this month. Google’s new speech models raise the opposite problem, generating a voice rather than reading one, and is Gemini 3.8 TTS safe covers its consent checks and watermarks. For how the bigger vendors document data handling, is Claude Haiku 5.5 safe reads Anthropic’s terms the same way. If the attraction of Falcon ASR is Arabic and multilingual support, multilingual AI support tools looks at that problem from the business side.
The verdict
Falcon ASR looks like a strong model for Arabic speech, and the demo is open about how its own page behaves. What it does not yet have is a published answer on retention and training at the server end. Treat it the way you would treat any hosted tool without a privacy notice: fine for test clips, not for anything you would not want kept.
Coursium is a mobile app that teaches people to use AI at work, and it is on the App Store. New models arrive weekly. Knowing where your data goes before you press upload, and checking what comes back, are the habits that carry over from one to the next. If that is what you want to practise, have a look at Coursium.
Frequently asked questions
Is Falcon ASR safe to use?
For test clips, yes. The public demo sends your audio to TII’s servers for transcription. The demo’s README documents that its own analytics hold no audio, transcripts or IP addresses and that its cache is cleared after 24 hours, but nothing published says how long TII’s API keeps audio or whether it is used for training. Keep sensitive recordings out until that is documented.
Can I download Falcon ASR and run it locally?
Not as of 9 October 2026. TII’s launch post points to a hosted demo and says API access and native apps are planned. We found no downloadable Falcon ASR weights on TII’s Hugging Face organisation.
Which languages does Falcon ASR support?
Arabic, including Modern Standard Arabic, Emirati and other Gulf dialects, plus English, French, Spanish and Portuguese, all with the same weights and no language setting, according to TII’s launch post of 7 October 2026.
How accurate is Falcon ASR?
TII reports an average word error rate of 20.92% across six public Arabic test sets and 5.74% across seven English ones. Those are the vendor’s own figures, and they still mean a transcript needs checking before you rely on it.