Voices API support arrived in OpenAI’s official Python library on 2 October 2026, when openai-python 3.24.0 added a method for creating custom voices from either a recorded audio sample or a plain text prompt. A day earlier, on 1 October, OpenAI made virtual Try On generally available in ChatGPT, so shoppers can see how clothes and accessories from shopping results might look on a photo of themselves. TestingCatalog put the two launches at the top of its Daily AI Brief on 3 October, and noted that the SDK change arrived “quietly”.

They look like unrelated releases, one for developers and one for shoppers. They share a theme, though. Both take something personal, your voice or your face and body, and turn it into a reusable reference that OpenAI stores and builds on with generative AI. One asks for a recorded statement of consent before it will clone a voice. The other asks for a selfie and keeps it for next time.

This article explains what changed in the Python SDK, how the Voices API works, the limits OpenAI places on it, how Try On works and where it is available, how it fits OpenAI’s shopping plans, and what developers, retailers and shoppers should do next.

What OpenAI Shipped on 1 and 2 October

openai voices api python sdk chatgpt virtual try on b tin can telephone pair joined by a string

The two launches arrived about 24 hours apart and through very different channels. Try On was announced in OpenAI’s product release notes and its help centre. The Voices API method appeared in a GitHub release with a one-line changelog entry and no blog post at all.

LaunchDateWhere it was announcedWho gets it
Virtual Try On and Favorites1 Oct 2026, marked GAChatGPT release notes and help centreChatGPT on mobile and web; no plan limit named
Voices API method in openai-python 3.24.02 Oct 2026, 15:54 UTCGitHub release and pull request 4013Organisations with custom voice access
Voices API voices designed from a text prompt2 Oct 2026API reference and SDK types onlyThe same organisations, in GPT-Live sessions only

The Voices API in the Python SDK

Release 3.24.0 of openai-python lists one new feature: “add custom voice creation and agent session events”. The pull request behind it, number 4013, was merged at 05:54 UTC on 2 October, and the release went out on GitHub about ten hours later. It adds a new Voices API resource, client.audio.voices, with a single create method that posts to the /audio/voices endpoint.

Virtual Try On in ChatGPT

OpenAI’s release notes for 1 October describe “New ways to shop in ChatGPT”. A “Try on” button now appears on product listings for clothes and accessories. Select it, take or upload a selfie, and ChatGPT Images generates a picture of you wearing the item. A second feature, Favorites, lets you save products to your ChatGPT Library.

Why the two belong in one story

Both features run on a stored likeness. The Voices API keeps a voice that later API calls can reuse by its ID. Try On keeps your reference photo for future try-ons. How each one handles consent and storage is the most useful thing to understand about them, and we come back to it below.

How the Voices API Works in openai-python 3.24.0

openai voices api python sdk chatgpt virtual try on c vocal booth with a door and a small window

The new Voices API code is generated from OpenAI’s API specification, like the rest of the library. It is short, but it says a lot about how OpenAI wants custom voices to be made.

One Voices API method, two ways to make a voice

client.audio.voices.create has two typed overloads. The first takes an audio_sample file, a consent recording ID and a name. The second takes a name, a text prompt and type="prompt". If you leave type out, the Voices API assumes you are sending an audio sample.

from openai import OpenAI

client = OpenAI()

voice = client.audio.voices.create(
    type="prompt",
    name="Warm narrator",
    prompt="A warm, calm narrator with a clear, measured delivery.",
)
print(voice.id, voice.type)

The example prompt is the one OpenAI uses in its own API reference. An async class, AsyncVoices, mirrors the same two overloads for asyncio code.

Cloning a voice through the Voices API

The audio-sample route is the one OpenAI documented first. You upload a short recording of the voice you want, alongside the ID of a separate consent recording made by the same speaker. Samples can be up to 10 MiB and must be MPEG, WAV, Ogg, AAC, FLAC, WebM or MP4 audio. The SDK sends the request as multipart form data.

Designing a voice with a Voices API text prompt

The Voices API prompt route is new in 3.24.0. Instead of a recording, you describe the voice in words and the service generates a synthetic voice to match. That is a natural language processing problem as much as an audio one, because the model has to turn adjectives such as “warm” and “measured” into pitch, pace and timbre.

The prompt can be up to 2,000 characters. An optional script_hint, also up to 2,000 characters, sets the words the voice speaks while it is being created; without one, a script is generated from your description. A model parameter accepts “auto”, the default, or the dated snapshot “2026-10-01”.

What the call returns

Every successful Voices API call returns a Voice object with five fields: id, created_at, name, object (always “audio.voice”) and type, which records whether the voice came from an audio sample or a prompt. The Voices API does not return preview audio, so you have to use the voice in a session to hear it. The table below compares the two Voices API modes field by field.

FieldVoice from an audio sampleVoice from a text prompt
nameRequired, 1 to 256 charactersRequired, 1 to 256 characters
type“audio_sample”, the default“prompt”, required
SourceAudio file, up to 10 MiBDescription, up to 2,000 characters
ConsentConsent recording ID, requiredNot used
Optional extrasNonescript_hint and model (“auto” or “2026-10-01”)
Where it can speakSpeech, Realtime, Chat Completions audio, GPT-LiveGPT-Live only

What "Quietly" Means for the Voices API

openai voices api python sdk chatgpt virtual try on d fitting room with a drawn curtain

TestingCatalog’s word was well chosen. OpenAI has published no announcement for the Voices API method, and its pieces arrived over about ten months.

The Voices API endpoint is older than the SDK method

Custom voices first appeared in OpenAI’s API changelog on 15 December 2025, when new audio snapshots shipped with “support for Custom voices for eligible customers”. In March 2026 the Python SDK learned to pass a custom voice ID to the speech, Realtime and Chat Completions endpoints. The specification bundled with version 3.23.0 already described the /audio/voices and /audio/voice_consents endpoints. What 3.24.0 adds is a typed Python method for the Voices API, plus the new prompt option.

The guide has not caught up

OpenAI’s custom voices guide still describes only the audio-sample route, and says voices “must be created through an API request”. The prompt option appears in the API reference and the SDK, but not in the guide. Anyone reading the guide alone would not know the Voices API can now design a voice from a description.

Voices API access is still gated

The guide is clear that “Custom voices are limited to eligible customers”, with a link to OpenAI’s sales team. An organisation can create at most 20 voices. A typed method in the SDK does not change who is allowed to call the Voices API, and the specification lists a 404 response for when “the endpoint is unavailable”.

There is no consent helper, and no delete call

The SDK’s audio folder now holds four resources: speech, transcriptions, translations and voices. There is no voice consents resource, even though the API offers create, list, retrieve, update and delete operations for consent recordings. Developers on the audio-sample route still upload consent with a raw HTTP request before they call the Voices API from Python. The Voices API itself has no call to list, retrieve or delete a voice, only to create one.

The Voices API release was also one of six openai-python versions in five days, which helps explain why it slipped past most coverage.

openai-python releases published per day, 28 September to 2 October 2026 (our count from the GitHub releases page)

28 Sep, version 3.20.0: 1
29 Sep, versions 3.21.0 and 3.22.0: 2
30 Sep, version 3.22.1: 1
1 Oct, version 3.23.0: 1
2 Oct, version 3.24.0 with the Voices API method: 1

The run spanned OpenAI DevDay on 29 September, when versions 3.21.0 and 3.22.0 added the GPT-6.1 Sol model identifier and computer use for beta agents.

openai voices api python sdk chatgpt virtual try on e open shoe box with its lid leaning on it

The most important constraints are not in the Python code. They sit in OpenAI’s guide and its API specification.

The Voices API consent recording

Before the Voices API will clone a voice, the speaker must record a fixed consent phrase. In English it reads: “I am the owner of this voice and I consent to OpenAI using this voice to create a synthetic voice model.” OpenAI lists the phrase in 16 languages, and warns that “any divergence from the script will lead to a failure”. The consent and the sample must come from the same person, and one consent can be reused for several attempts by the same speaker.

Recording rules for Voices API samples

For GPT-Live, the sample needs at least five seconds of actual speech and at least 15 transcribed text tokens, and OpenAI recommends 10 to 30 seconds with several complete sentences. The general guide caps samples at 30 seconds. It also advises a quiet room, a professional XLR microphone and a steady 7 to 8 inches from the mic, because “the model copies exactly what you give it”.

Voices API prompt voices only work in Live

This is the catch for the new Voices API option. The specification says voices created from text prompts “are supported only in Live, not in Realtime or the speech endpoint.” OpenAI’s speech-to-speech GPT-Live API, which we covered when GPT-Live-1 reached the API in September, is the only place a prompt-made voice can speak. Even there, the guide says gpt-live-1 supports custom voices with English accents, and the voice is fixed when a session starts.

Voices API keys, scopes and failures

The guide asks for a project-scoped API key approved for both GPT-Live and custom voice creation through the Voices API. Reading consent phrases and using a voice needs the api.voices.read scope, while creating consents and voices needs api.voices.write. A deleted or revoked voice, or a consent from another project, can surface as a 404. Browser recorders that label audio as “audio/webm;codecs=opus” are rejected unless the upload uses the base “audio/webm” type. Here is where each kind of Voices API voice can be used.

Where you use the voiceVoice from an audio sampleVoice from a text prompt
Text to speech (/audio/speech)YesNo
Realtime APIYesNo
Chat Completions with audio outputYesNo
GPT-Live sessionsYes, with English accentsYes, with English accents

The wider market is moving fast on synthetic speech. ElevenLabs launched Eleven v4 and v4 Turbo only days earlier, so the Voices API prompt option arrives in a crowded field.

The Other Changes in openai-python 3.24.0

openai voices api python sdk chatgpt virtual try on f card catalogue cabinet with one drawer pulled out

The Voices API was the headline, but the same pull request touched 58 files. Three other changes matter to developers.

Agent session webhooks

The SDK now recognises five webhook events: agent.session.created, agent.session.in_progress, agent.session.idle, agent.session.failed and agent.session.action_required. They let a server react when a hosted agent session starts, goes idle, fails or needs a human decision, instead of polling for status.

Ultrafast removed from beta agents

The “ultrafast” value was removed from the service tier options when creating or updating beta agents and sessions. OpenAI launched Ultrafast on 29 September as a premium speed tier for GPT-6 Astra. The pull request does not say why the value was withdrawn from these calls.

Usage and image clean-up

Organisation usage results gain an input_cache_write_12h_tokens field. The image parameters now describe dall-e-2 and dall-e-3 as retired on 12 May 2026, and list chatgpt-image-latest alongside the GPT image models.

How ChatGPT's Virtual Try On Works

Try On is far simpler than the Voices API, and it is aimed at a much larger audience. OpenAI’s help article sets out the flow in three steps.

Three steps to a try-on image

First, select the “Try on” button on a product listing for clothes or accessories. Second, take or upload a selfie. Third, view the generated image. You can also upload a picture of clothing or accessories in a conversation and ask ChatGPT to show how the item could look on you.

Bring your own item

You are not limited to products ChatGPT suggests. Android Authority’s reviewer, testing in Canada, shared a link to an overcoat and saw it on their own photo. They also uploaded a picture of an outfit, asked for similar products to buy, and got both the products and a try-on. TechCrunch notes you can describe a style and ask ChatGPT to shop for the pieces, or upload photos of celebrity outfits and ask it to find the pieces they are wearing.

Favorites and folders in the Library

To save a product, you select its bookmark icon and it goes to Favorites. You can also create folders to organise your finds. Saved items live in the ChatGPT Library, and TechCrunch reports they sit alongside your try-on images.

Powered by ChatGPT Images 2.5

OpenAI says Try On uses ChatGPT Images 2.5, the image model it launched on 8 September 2026. OpenAI’s launch post says the model “is better at preserving the subjects in your reference photos” and cuts generation latency by up to 50% compared with Images 2.0. It also says people create more than 3 billion images a week across ChatGPT Images and its image models in the API. Virtual try-on has been a hard computer vision problem for years, and keeping a person recognisable while changing their clothes is exactly the subject-preservation skill OpenAI claims for the new model.

Where Virtual Try On Is Available

OpenAI’s description is short. The release notes mark the feature “GA” and say it is “Available in ChatGPT on mobile and web.”

A global launch on mobile and web

TechCrunch reported that OpenAI announced “the global launch” of both shopping features. The release notes name no plan restriction, which suggests free users get it too, though OpenAI has not said so directly. Android Authority confirmed that both features were live for its reviewer in Canada on mobile and web.

What the reports disagree on

OpenAI’s help article and release notes mention only a selfie. Engadget and Android Authority both describe being asked for a full-body photo as well, after the selfie. The app appears to ask for both on first use, even though the written instructions mention one. Expect to need a full-length photo for a convincing result.

Try On and OpenAI's Shopping Strategy

Try On is the latest step in a shopping effort that has already changed direction once this year.

From Instant Checkout to product discovery

OpenAI launched Instant Checkout in September 2025, letting people buy some products without leaving ChatGPT. On 24 March 2026 it changed course. “We’ve found that the initial version of Instant Checkout did not offer the level of flexibility that we aspire to provide,” it said, and it moved its focus to product discovery. Engadget now describes the feature as discontinued and says OpenAI takes no cut of sales. Oddly, the shopping help article still says an Instant Checkout option may appear “for some eligible products and merchants”.

Results are not ads

The help article states that “Product results are selected independently by ChatGPT and are not ads, nor influenced by any OpenAI partnerships.” Ads in ChatGPT are separate. Products are chosen from structured merchant data and third-party content, and Shopify merchants are already included through Shopify Catalog.

The competition

Google launched its own virtual try-on in July 2025, more than a year before OpenAI’s, a point TechCrunch was quick to make. Agent start-up Instinct began pushing product recommendations days earlier, to a mixed reaction. Retailers are also pushing back against shopping agents, as when Amazon barred Meta’s Muse AI from shopping on its site. On the merchant side, Shopify’s Canvas lets sellers build a store by chatting with AI.

DateChatGPT shopping milestone
April 2025Basic shopping results arrive in ChatGPT search
September 2025Instant Checkout launches
24 March 2026OpenAI steps back from Instant Checkout and focuses on visual product discovery
8 September 2026ChatGPT Images 2.5 launches
1 October 2026Try On and Favorites become generally available

The Privacy Question Both Launches Raise

Putting the Voices API and Try On side by side shows how differently OpenAI treats a voice and a face.

Your photo is kept

Try On saves your reference photo so you do not have to upload it again. You can change or delete it under Settings, then Personalization, then Reference photos. That is a sensible control, but it is opt-out: the photo stays until you remove it.

Training defaults

Engadget points out that, by default, images uploaded from a personal ChatGPT account can be used for training unless you opt out in your data controls. A selfie for a try-on is no different from any other upload in that respect.

Voices API consent versus a selfie

The Voices API will not clone a voice until the speaker reads a fixed consent statement, and the recording must match the sample. Try On has no equivalent step. OpenAI’s help article does not describe any check that the selfie shows the person using the account. The stakes are lower for a jacket than for a cloned voice, but it is a gap worth knowing about.

The UK view

Under UK GDPR, a photo of your face or a recording of your voice is personal data. The ICO’s guidance explains that it becomes special category biometric data when it is processed to recognise a specific person. Courts are also taking voices seriously: a Japanese court has ruled that human voices are protected in a landmark AI case. Businesses that build on the Voices API or send customers to Try On should review their data protection obligations first.

QuestionVoices APIChatGPT Try On
What is storedA voice, reusable by its IDYour reference photo
Consent stepFixed spoken phrase from the speaker, for cloned voicesNone described beyond the upload
Who can use itEligible API customersChatGPT users on mobile and web
Limits20 voices per organisation; samples up to 30 secondsClothes and accessories only
How to remove itConsent recordings can be deleted; the spec has no voice delete callSettings, Personalization, Reference photos

What Developers, Retailers and Shoppers Should Do Now

The two launches call for different actions, depending on which side of them you are on.

Developers

Upgrade to openai-python 3.24.0 or later to get the typed Voices API method, but confirm custom voice access with OpenAI before you plan around it. Build the Voices API consent step into your own product, so each speaker records the phrase and you keep the consent ID. Do not promise Voices API prompt voices anywhere except GPT-Live. If you run hosted agents, wire up the new session webhooks. Our software development team can help you plan an integration.

Retailers

Try On works on the product images and data ChatGPT already holds, so the quality of that data now affects how your clothes look on a shopper. Shopify merchants are included through Shopify Catalog, and other merchants can apply to send OpenAI a direct product feed. Clear measurements and a fair returns policy matter more than ever, because OpenAI’s help article warns that try-ons “do not guarantee fit or size”.

Shoppers

Treat a try-on image as a rough preview, not a fitting. Check the merchant’s size chart and returns policy before you buy. If you would rather OpenAI did not keep your photo, delete it from Reference photos after use, and review your data controls.

Voices API and Try On FAQ

What is the OpenAI Voices API?

The Voices API is OpenAI’s endpoint for creating custom voices. It can clone a voice from an audio sample backed by a consent recording, or design a new synthetic voice from a text description. The Python SDK added a method for it in version 3.24.0.

Which version of the Python SDK adds the Voices API?

The Voices API method arrived in version 3.24.0 of openai-python, released on 2 October 2026. The method is client.audio.voices.create, with an async version in AsyncVoices.

Can every developer use the Voices API?

No. OpenAI limits custom voices, and so the Voices API, to eligible customers, and organisations need to contact its sales team for access. Each organisation can create up to 20 voices.

Do Voices API prompt voices work with text to speech?

No. Voices designed from a text prompt work only in GPT-Live sessions. Voices cloned through the Voices API from an audio sample also work with text to speech, the Realtime API and Chat Completions audio output.

Is ChatGPT Try On free?

OpenAI’s release notes say Try On is available in ChatGPT on mobile and web and name no plan restriction. OpenAI has not published a plan-by-plan list, so check the app if you are on the free tier.

How do I delete my Try On photo?

Open Settings, then Personalization, then Reference photos. Select the photo and choose Change photo to replace it, or the delete icon to remove it.

Does Try On guarantee the right size?

No. OpenAI says try-on images “may not represent the product or your appearance exactly and do not guarantee fit or size”. Check the merchant’s measurements, product details and returns policy before buying.

References and Further Reading