Short answer: use an external speech-to-text service, then send the transcript to a multi-model gateway for summarization. A single key sounds cleaner, but it is the wrong selection criterion when the audio capability itself is not available; the useful boundary is “transcript text in, evaluated...
Source: [Dev.to](https://dev.to/jasperflint6947/speech-to-text-plus-multi-model-transcript-summaries-one-api-key-or-two-1ei5)