Discontinuation of Sora API, arrival of Hy Image 3.5, and Google’s video and audio updates
Period: September 20 – September 26, 2026
We have summarized the major trends in image, video, and audio generation models and related services. We look back at not only notable new features but also important milestones such as the discontinuation of existing services.
September 22 | Hy Image 3.5 preview: An image model designed with post-generation “retouching” in mind
The “Hy Image 3.5 preview” announced by Tencent is a model that goes beyond text-to-image generation and also excels at partial editing based on reference images.
According to the official announcement, the accuracy of text rendering, consistency in composition and layout, and the ability to maintain the characteristics of the original image have been significantly improved.
It can ingest up to five reference images and output at a maximum resolution of 2K. It has already been integrated into the company’s collaboration tool, “WorkBuddy,” which offers a workflow where users can convey requests in a chat format or upload images directly to instruct modifications.
For teams that want to create variations while maintaining the text and product details within a poster, this seems to be a more practical option than simple one-shot generative AI.
However, it remains to be seen how well the improvements in image quality and consistency touted by the official announcement will satisfy Japanese font expression and complex design requirements.
When introducing it into actual work, it is necessary to verify with real project data whether fine logos are being blurred or unintended alterations are occurring.
Note that usage on WorkBuddy consumes points, so even though it is a preview version, it is not unconditionally unlimited.
September 23 | Google Vids AI video generation opened to personal accounts
For those who have created video materials with generative AI but found the task of moving them to timeline editing software to arrange them tedious, this update to Google Vids is good news.
If you have a Google account, you can now try video generation using Gemini Omni 1.1 Flash for free by simply opening Vids from a PC browser and selecting “Create AI videos.”
While Omni 1.1 Flash itself is an existing model, the topic here is that it has been seamlessly integrated as a production environment for individuals.
Vids, which previously had a strong enterprise focus, now allows users to proceed from material generation to timeline editing and final export entirely within the browser.
Production features have also been updated with a focus on practicality.
In addition to 1080p scene generation, it now supports upscaling of existing clips.
It also includes an “Extend” feature that generates continuations while maintaining consistency in lighting, subjects, and backgrounds, as well as the ability to specify the duration of generated clips in detail.
This should be useful in situations where you want to adjust cuts to match the duration of a narration.
However, just because it is free does not mean it can be used without limits.
Google provides generous generation quotas for paid AI plans and has prepared management features and additional quotas for Workspace Business/Enterprise users.
Since the specific generation limit for free accounts is not clearly stated, it is safer to check the display on the management screen before starting work.
Also, an invisible digital watermark, “SynthID,” is embedded in the output video frames.
When combining your own photos, client logos, or background music with professional videos, you must manage the usage rights for each separately.
The fact that the tool itself has been opened for free and the commercial use of all materials being cleared are separate issues.
Regarding the narration feature “Gemini 3.8 Flash-Lite TTS,” which supports over 100 languages, the Vids announcement states it is “coming soon.”
The release for the TTS side published on the same day contains descriptions that could be interpreted as a sequential rollout, so it is important to note that it has not been reflected on all accounts immediately.
September 23 | Gemini 3.8 TTS: Fine-tuning voice tone and performance
Google simultaneously announced two voice generation models: “Gemini 3.8 Flash TTS” and “Gemini 3.8 Flash-Lite TTS.”
The former focuses on the expressiveness of character voices and long-form content, while the latter is intended for mass-producing dubbing and audio guides at a low cost.
A key feature is that you can not only choose voices from presets but also design your own preferred voice from scratch by instructing the role, voice quality, and accent in the prompt.
The higher-end Flash TTS model supports over 100 languages and dialects and includes over 2,000 preset voices.
Furthermore, if you have the rights to use a voice, you can create a reusable voice profile from just a 30-second sample.
To prevent misuse, a step has been incorporated to verify consent audio from the owner of the sample voice.
On the script side, you can control speech speed, emotional inflection, and dialect switching on a line-by-line basis.
In addition to embedding non-verbal expressions like laughter, sighs, and backchanneling, it is also possible to have two characters perform a conversation from a single script, making it feasible to output long-form audio like talk podcasts or audiobooks.
It can be said that this has moved beyond the realm of traditional “text-to-speech” and closer to the work of a director giving acting guidance to voice actors.
The reproduction of Japanese is also subject to verification.
According to Google’s announcement, it recorded an overall score of 71.4 and an accent evaluation of 60.8 on Hume AI’s “Voice Design Benchmark,” and also ranked highly in multiple languages, including Japanese, in the blind-test format “Voice Arena.”
However, benchmark scores do not necessarily translate directly to your own content.
Intonation of proper nouns, the “pauses” between lines, and fluctuations in voice quality during long-form playback need to be compared by actually listening through your own company’s scripts.
For developers, availability has begun via the Gemini API and Google AI Studio under the model IDs gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts.
Integration into Gemini Notebook and Google Vids has been announced, but API availability for Gemini Enterprise is currently in “coming soon” status.
The standard API pricing until December 31, 2026, is $0.50 per 1 million tokens for text input, and $9 for Flash and $6 for Flash-Lite for audio output (approximately $0.00225 and $0.0015 per 10 seconds, respectively).
Starting January 1, 2027, prices for both input and output are scheduled to double. Since the terms of service state that data sent via the free-tier API is used to improve Google’s services, using a paid plan where data is not used for training is a prerequisite when handling customer audio data or confidential manuscripts.
Generated audio is watermarked with SynthID, and support for C2PA provenance information is also underway.
However, these protection features do not automatically handle usage licensing on your behalf.
Note that the audio cloning feature in Google AI Studio is restricted in certain U.S. states (Illinois, Texas), the EEA, the U.K., Switzerland, and India due to legal regulations.
The ‘Voice remixing’ feature, which allows for post-processing the voice quality and pitch of existing audio, remains at the announcement stage.
September 24 | Sora API Discontinued. The standard for ‘video generation’ that pioneered and elevated video generation models.
Since the first Sora demo video was released in February 2024, the public’s perspective on generative video AI has changed fundamentally.
It went beyond the level of still images shaking slightly; the depth of the space remained intact even when the camera panned dynamically, and subjects maintained their appearance even after moving out of frame and reappearing.
Including a maximum duration of one minute, Sora showed the world the potential for AI to simulate ‘space itself accompanied by the passage of time.’
Although the models at the time still had unnatural physics and distortions in long clips, they redefined the very criteria by which creators evaluate video AI.
It wasn’t just about the beauty of each frame, but whether a series of actions could be completed as the same person, and whether the background perspective was consistent.
In other words, the achievement of creating a culture that asks ‘whether it works as a video production’ remains unshaken.
The December 2024 commercial version supported extending the duration of existing materials, partial corrections, compositing features, and even specifying cuts via storyboards.
It stepped into professional workflowsânot just outputting a single video and finishing, but refining the composition, adjusting cuts, and requesting retakes if necessary.
The following year’s ‘Sora 2’ continued to evolve with video synchronization for dialogue and sound effects, and more realistic physical behavior. If looking toward film and advertising production, Sora continued to embody as features the standard that one must be able to handle not just image quality, but acting, continuity, sound, and direction in a one-stop manner.
The web and app versions of Sora ended their service on April 26, 2026, and on September 24, the provision of the API was officially terminated.
While Sora as a product has come to an end, the quality standards for video generation that they pushed to the limit will surely remain as a baseline for evaluating all competing tools that appear in the future.
September 24 | Gemini 3.8 Live gets a ‘face’: Official launch of interactive avatars.
Following the ‘Gemini 3.8 Live’ that can operate external tools while continuing a conversation, which we covered last week, this week ‘Live Avatar’ that visualizes that interaction was officially released for Gemini Enterprise.
It receives the user’s voice and camera footage, and moves the avatar’s mouth and expressions in real-time to match the voice of the AI’s response.
The model ID for developers is gemini-3.8-live.
It accepts 16kHz audio, 1 frame-per-second video, and text as input, and outputs 24kHz audio and 24fps MP4 video in real-time.
The ability to output lip-synced video directly from the foundation model without inserting a third-party avatar rendering tool is a major architectural advancement.
It is possible to not only use preset avatars but also to create custom avatars from a single portrait photo.
However, custom creation requires pre-approval for enterprise use and an identity verification process, so it is not designed to allow just anyone to animate someone else’s photo without permission.
It supports 97 languages, automatically follows along even if the language switches during a conversation, and reproduces mouth movements matched to the pronunciation.
As a suggested use case, a demo was introduced for a hotel check-in counter where the AI converses with a guest while checking and updating the reservation system in the background.
Since it also supports real-time recognition of camera footage and screen sharing, applications such as customer service windows where you receive consultations while showing documents are also realistic.
However, quantitative data regarding the naturalness of Japanese-specific subtle honorific expressions and pronunciation is not included in this release.
The platform is limited to Gemini Enterprise, and connection endpoints are only in the U.S. and the EU.
It has not trickled down to Gemini Live for general users, and there is no mention of closed-network processing within the Japanese domestic region.
Although SynthID is applied to the output video and audio, there are many issues left for the operating company to resolve when incorporating it into customer service terminals, such as clearly stating that it is an AI response, clearing portrait rights, and establishing conversation log retention policies.
Summary
While the Sora API, which made history in video generation, quietly exits, Google Vids has lowered the hurdle for video production for individuals, and Tencent has launched Hy Image3.5, which makes generation and editing a continuous process.
In the audio field, Gemini TTS has stepped into ‘voice performance,’ and with the arrival of Live Avatar, real-time expressions have been added to voice interaction models.
The choices for expression are expanding rapidly, but the usage environment, billing structure, data handling terms, and service lifespan vary for each model.
Before immersing yourself in the creative work of extending videos, refining images, and adding voices, you must determine where and how to store the exported results and how to guarantee rights.
You need to determine the exit strategy for the entire workflow in advance.
