OpenAI Has Added Audio File Uploads to ChatGPT
Paid users can now analyze audio files up to 512MB, expanding the platform beyond text and image inputs.
Updated on Oct. 7, 2026 in Artificial Intelligence

Live Poll
Would you use an AI tool to transcribe and summarize your recorded meetings?
OpenAI has officially released audio-file upload functionality for paid ChatGPT subscribers and enterprise workspaces. The feature allows users to process various audio formats, though video files remain unsupported.
Why it matters
This update broadens the multimodal capabilities of ChatGPT, enabling direct analysis of recorded speech and sound without requiring separate transcription tools. It marks a shift toward integrating audio-centric workflows directly into the LLM interface.
The system supports a maximum file size of 512MB for formats including WAV, MP3, OGG, FLAC, AAC, M4A, PCM, WebM, and MP4. While the model processes these inputs, users may encounter processing timeouts on exceptionally long recordings.
The players
OpenAI
An AI research and deployment company known for its GPT large language models and the ChatGPT interface.
The details
To utilize the function, users attach a supported audio file to their session and define the required task. OpenAI utilizes its underlying architecture to parse these files, though the company cautions that transcripts may contain errors and speaker identification may remain unreliable. For files processed through the Data Analysis tool, the system may break long audio into segments for systematic evaluation.
Timeline
September 29, 2026: OpenAI announced the Meetings plugin at DevDay.
October 6, 2026: OpenAI released the audio-file upload feature for ChatGPT.
The Tech Race
This release builds on the trajectory established by the Meetings plugin announced at DevDay. It positions OpenAI to compete more directly with integrated audio-transcription and analysis services by bringing these capabilities into the primary chat interface.
Paid users and enterprise workspaces can immediately utilize this feature to analyze audio content, while free plan users remain excluded from the tool. Users should note that while most common audio formats work, video files will not be accepted by the system.
The takeaway
The move underscores a push for native multimodal intelligence that bypasses traditional secondary transcription workflows. Users should monitor for future updates regarding processing timeouts and improved speaker identification accuracy for complex, multi-person recordings.
Further reading
For broader trends in LLM capabilities, explore our Artificial Intelligence section.
Live Poll
Would you use an AI tool to transcribe and summarize your recorded meetings?







