A buyer-focused log of material changes discovered across Chinese AI reporting and verified against vendor evidence. Publication dates and vendor source dates stay separate so older announcements are never presented as new.
Primary-source verified; affected models, API surfaces, regions, error rates and root cause were not disclosedOperational availability and incident transparency
MiniMax records repeated short LLM error incidents across September 29 and 30
MiniMax's official status page records twelve separate periods of elevated errors for its Large Language Models component on September 29 between 00:03 and 10:04 UTC. Each was marked resolved after one to eleven minutes. The page records another LLM incident on September 30 from 01:04 to 01:10 UTC and currently reports all systems operational. The incident notices do not identify specific models, API endpoints, accounts, regions, error rates or a root cause.
Buyer impact
Teams using MiniMax LLMs in production should treat the pattern as evidence of brief, recurring service instability rather than a continuing outage. Review failures in your own account for the cited UTC windows, keep retries bounded and idempotent, use circuit breakers and a tested fallback route where the workload requires it, and prevent retry logic from duplicating billable work. Do not infer that every MiniMax model, region or API was affected, or that an SLA was breached, because the public notices do not provide that scope.
Vendor source date
2026-09-29 to 2026-09-30 (official incident posts); status page accessed 2026-09-30
Primary-source verified with billing and endpoint-documentation gapsPreview multimodal agent API availability and interface compatibility
Meituan adds LongCat-2.5-Preview to its OpenAI- and Anthropic-compatible API
LongCat's official September 25 change log says LongCat-2.5-Preview is now available with image understanding, coding and integrations for mainstream coding environments. The current quick-start guide lists the model on both OpenAI- and Anthropic-compatible endpoints with a one-million-token context window and maximum output of 128,000 tokens. The documentation is not yet fully aligned: the detailed chat-completions reference still names LongCat-2.0 and text-only input, while the pricing page says pay-as-you-go currently supports LongCat-2.0 even though the FAQ says token packs support LongCat-2.5-Preview.
Buyer impact
Teams evaluating an agentic coding or multimodal model can test LongCat-2.5-Preview through familiar OpenAI or Anthropic client formats instead of adopting a proprietary transport. Treat this as preview availability, not protocol parity or production readiness. Confirm the live model list and entitlement in the account, test the exact image-input schema and tool behaviour, cap context and output deliberately, and handle per-key 429 responses. Do not assume that LongCat-2.0 pay-as-you-go prices apply: verify the available billing method, current unit price, token-pack terms, quota, region and data path, taxes, service levels and provider terms before procurement or a production cutover.
Primary-source verifiedPrivate-beta multi-model gateway API capabilities
Tencent Cloud extends its model-router API for video, reranking and decision models
Tencent Cloud's September 25 API change log adds VideoConfig, RerankConfig and DecisionsConfig to the create and modify operations for Cloud Model Router, adds Capabilities to model-association responses, and extends router details with those configurations plus a load-balancer identifier. Tencent describes Cloud Model Router as a unified model gateway, but the current product documentation labels the service as private beta.
Buyer impact
Teams evaluating a customer-operated multi-model gateway can now assess video-generation, reranking and decision-model routes through the same documented router API instead of treating it as text-model-only infrastructure. This is an interface change, not proof that every upstream model, region or account is enabled. Pin the API version, test create, modify and describe responses, and verify configuration defaults, upstream credentials, routing and failure behaviour, logging, data paths, load-balancer ownership, quota and billing. Confirm private-beta access, supported regions and models, provider terms, production support and service levels with Tencent before committing a production migration.
Vendor source date
2026-09-25 (API change log); product and API documentation accessed 2026-09-26
Primary-source verifiedHosted DeepSeek peak and off-peak billing schedule
Tencent Cloud makes weekends off-peak for selected hosted DeepSeek routes
Tencent Cloud's official pricing-adjustment notice says that, from 00:00 Beijing time on September 26, weekends are entirely off-peak for selected non-direct-supply DeepSeek models in TokenHub and the Agent Development Platform. Weekday peak windows remain 09:00-12:00 and 14:00-18:00 Beijing time. The notice identifies TokenHub routes including deepseek-v4-flash-0731, deepseek-v4-pro-0813 and deepseek-v4.1-flash, with affected Flash and Pro routes also listed for the Agent Development Platform.
Buyer impact
Buyers using the named Tencent-hosted routes can schedule flexible batch evaluation and generation for the Beijing-time weekend to avoid a peak window. Do not apply this rule to DeepSeek's direct API, Tencent's vendor-direct routes or unlisted models: each can have a different schedule, unit price, service agreement and support model. Verify the exact model ID and supply type shown in the account console, current peak and off-peak rates, billing timezone, credits, quotas, taxes, data location, provider terms and production support before changing workloads or forecasting spend.
Primary-source verifiedRegional preview decision-model API availability and billing
Alibaba Cloud adds a preview decision model in Beijing and Singapore
Alibaba Cloud Model Studio's lifecycle page lists decision-model-preview as released on September 24 in China (Beijing) and Singapore. The provider describes the preview model as supporting parallel classification, yes-or-no decisions and scoring, returning probability and confidence values. The current pricing documentation says usage is metered on input tokens but is temporarily free during the stated limited-time period.
Buyer impact
Teams building policy, routing or triage workflows can now test a purpose-built decision endpoint in two documented regions instead of forcing the task through a general chat model. Treat the model as preview and the free period as temporary: instrument input-token volume and decision quality, define confidence thresholds and fallbacks, and budget for a future paid rate that has not been established here. Confirm the live model ID and endpoint, workspace and API-key region, account entitlement, quota, current price and free-period terms, output stability, data location, provider terms and production support before using it in a consequential workflow.
Vendor source date
2026-09-24 (official model lifecycle entry); pricing and lifecycle documentation accessed 2026-09-26
Primary-source verified with unresolved official documentation conflictMultimodal API model availability, alias routing and billing clarification
DeepSeek keeps V4 Pro available after reversing the September 14 routing plan
DeepSeek's current English API change log, Quick Start, pricing page and Anthropic compatibility guide now say that V4 Pro remains available after September 14 with billing unchanged. The pricing page lists deepseek-v4-pro as DeepSeek-V4-Pro-0813 at USD 0.022/0.044 per million cache-hit input tokens, USD 0.66/1.32 per million cache-miss input tokens and USD 1.98/3.96 per million output tokens at off-peak/peak rates, with concurrency 500. This supersedes our earlier Signal that treated the announced V4.1 Flash routing as effective. However, DeepSeek's English launch article and Chinese pricing page still say deepseek-v4-pro requests route to V4.1 Flash at Flash rates, so the provider's official documentation remains internally inconsistent.
Buyer impact
Direct DeepSeek API buyers should not assume either V4 Pro continuity or V4.1 Flash alias routing from one official page alone. Regression-test the live endpoint, record the resolved model and behaviour in telemetry, and reconcile account-console usage and invoices before treating the current V4 Pro rates as procurement facts. Teams using Anthropic-compatible model names should also verify mapping because the current guide maps Claude Opus aliases to deepseek-v4-pro. Do not extrapolate DeepSeek's direct-API behaviour, prices or limits to Tencent Cloud, Alibaba Cloud or another host. Confirm account entitlement, endpoint, regional availability, quotas, data handling, taxes, provider terms and support before migration or committed spend.
Vendor source date
2026-09-10 (original release); current English API change log, quick start, pricing and Anthropic compatibility documentation accessed 2026-09-18
Primary-source verifiedReal-time avatar and live-stream editing APIs
Vidu adds S2 real-time avatar and live-stream editing APIs
Vidu's September 15 platform update adds S2-Avatar for real-time interactive avatars and S2-Editing for live-stream video editing. The current model map and API documentation describe WebSocket session control plus RTC media transport, with avatar component, offline and real-time integration modes and editing tasks for style transfer, virtual try-on, subject editing and background replacement. The component-based avatar and stream-editing documents bill successful sessions at one credit per second, stated as CNY 0.03125 per second, while the official API repository identifies separate global and mainland-China hosts.
Buyer impact
Buyers evaluating Vidu S2 need a streaming architecture rather than the request-and-poll flow used by asynchronous video generation. Plan for server-side API-key proxying, WebSocket heartbeats and teardown, supported RTC providers, moderation, idempotent session creation, session concurrency and per-second cost controls. Billing starts after the documented successful connection acknowledgement, and third-party RTC services can add separate cost and data-path obligations. Verify account entitlement, the correct global or mainland-China host and key, regional availability, quotas, moderation behaviour, taxes, data handling, provider terms and production support before deployment.
Vendor source date
2026-09-15 (official platform update); official model, API and billing documentation accessed 2026-09-18
Primary-source verifiedManaged coding model alias, thinking controls and context availability
Kimi Code upgrades the kimi-for-coding alias to K2.8 Preview
Kimi Code has fully rolled out K2.8 Preview behind the existing kimi-for-coding model ID. The official documentation says clients and third-party tools do not need a configuration change, and now maps the alias to K2.8 Preview with low, high and max thinking-effort levels; max is the default when no effort is supplied. The release also advertises up to a one-million-token context window across membership tiers. Turning thinking off routes K3-series and K2.8 Preview requests to K2.8 Preview without thinking. The performance and efficiency descriptions are vendor claims; this announcement applies to the Kimi Code managed membership service and does not establish a matching model ID, price or rollout on Moonshot's separate pay-as-you-go API.
Buyer impact
Existing Kimi Code integrations that call kimi-for-coding can receive different model behaviour without an identifier change. Regression-test representative repositories and agent workflows for code quality, tool use, latency, token consumption, context handling and accepted-result quality; record the observation date and effective model in deployment telemetry where possible. Review effort mappings in third-party clients: an omitted value uses the model default, while unsupported values can fail. Do not assume that advertising a one-million-token window means every client configuration should send that much context or that quota consumption is unchanged. Confirm the live Kimi Code endpoint, membership entitlement, current credit and rolling limits, account and regional availability, terms, data handling and production support before procurement. Do not infer availability or pricing on platform.moonshot.ai from this Kimi Code update.
Vendor source date
2026-09-11 (official Kimi Code rollout; documentation accessed 2026-09-14)
Primary-source verifiedRegional high-speed inference mode and API pricing
Alibaba Cloud adds qwen3.8-max-prime for higher-throughput inference in Beijing
Alibaba Cloud Model Studio's current Chinese Prime-mode documentation lists qwen3.8-max-prime in China (Beijing). The mode is documented as providing roughly 1.5 to 2 times the output TPS of the standard API while retaining the original model's capabilities and usage limits. The current price table lists CNY 24 per million input tokens, CNY 72 per million output tokens and CNY 3 per million cache-hit input tokens, with no free quota. The standard qwen3.8-max row remains CNY 12 input and CNY 36 output with a one-million-token free quota, so the Prime route carries a two-times headline input and output unit price. Alibaba's English Prime page, accessed on September 10, does not yet list qwen3.8-max-prime.
Buyer impact
Teams with latency-sensitive coding, agent or real-time workloads can now run an account-level comparison between the standard and Prime qwen3.8-max routes in Beijing instead of assuming that more throughput requires a different base model. Treat the documented 1.5-to-2-times TPS as a service-mode description, not a workload guarantee: measure time to first token, sustained output rate, end-to-end latency, concurrency, cache behaviour and accepted-output quality with representative requests. The published Prime listing is Beijing-only for this model and uses a workspace-specific endpoint; do not infer Singapore or global availability from the standard model's regional footprint. Confirm live console access, exact model ID, endpoint and API-key region, current price, quota, taxes, data location, service levels and the unresolved Chinese-English documentation difference before procurement or production cutover.
Vendor source date
2026-09-09 13:41:51 (Prime documentation update; pricing and English documentation accessed 2026-09-10)
Primary-source verifiedCustomer-operated multi-provider video API gateway
Tencent Cloud documents customer-operated Kling and Vidu video API routing
Tencent Cloud's current AI Gateway documentation describes a customer-operated route for video generation across Kling and Vidu. The gateway can hold provider credentials, expose video task creation and retrieval routes, and connect to services using Kling OpenAPI or Vidu request protocols. Tencent's product log records the video-model integration as an August update; the detailed implementation guide was updated on September 7.
Buyer impact
Teams that already contract with Kling or Vidu can evaluate a gateway inside their own cloud account instead of giving a service intermediary their provider keys. This changes the data-path decision: prompts, media, credentials and task metadata pass through the customer's Tencent Cloud gateway, so its region, logging, retention, access controls and failure behaviour must be reviewed alongside each upstream provider. The documentation does not publish gateway pricing, supported regions, service levels, provider-plan eligibility or protocol parity. Verify the live console, account entitlement, exact request fields, task polling, fallback behaviour, quotas, data location and total cost before treating the route as production-ready.
Primary-source verifiedRegional cloud API transport security requirement
Tencent Cloud sets a December 1 TLS 1.2 minimum for APIs outside mainland China
Tencent Cloud's official notice says that, from 00:00:00 Beijing time on December 1, its APIs in listed regions outside mainland China will stop accepting TLS 1.0 and 1.1. Clients and applications must use TLS 1.2 or later. The affected list includes Hong Kong, Taipei, Singapore, Tokyo, Osaka, Seoul, Bangkok, Jakarta, Johor Bahru, Silicon Valley, Virginia, Frankfurt, Sao Paulo and Riyadh. Tencent provides an account audit path for identifying older-protocol calls.
Buyer impact
Teams using Tencent Cloud as part of a China AI deployment should inventory control-plane and SDK calls in the listed regions, identify obsolete runtimes or TLS libraries, and test TLS 1.2-or-later connectivity before December 1. The announcement applies to Tencent Cloud APIs generally; it does not explicitly confirm whether every product-specific model-inference hostname, including TokenHub endpoints, is in scope. Confirm the exact hostname and regional endpoint with Tencent, preserve current account evidence, and test both management and inference paths rather than assuming that a modern SDK alone removes every dependency.
Primary-source verifiedLegacy AI video API retirement and immediate service shutdown
Tencent retires four legacy Hunyuan video APIs and stops new sales for two more
Tencent Cloud's official notice lists six legacy Hunyuan specialist video interfaces for phased retirement. At 00:00 on September 3, ImageGameToVideo, VideoVoice, VideoStylization and VideoEdit stopped new sales and their API services were shut down. ImageAnimate and PortraitSing also stopped new sales at that time, but their APIs remain scheduled to shut down at 00:00 on September 30. Tencent directs customers toward TokenHub options including Hy-Video-v1.5, Youtu-Video-HumanActor, Vidu-Video-q3.0-pro and Kling-Video-V3.
Buyer impact
Teams that still depend on any of the first four interfaces should treat the retirement as an active production dependency issue, identify failed or residual calls and move traffic only after a controlled replacement test. Users of ImageAnimate or PortraitSing have a remaining migration window before September 30, but cannot make new purchases through those legacy services. Do not assume any recommended TokenHub model is a drop-in replacement: Tencent's visual migration guide says model identifiers, call addresses, request parameters and billing can change. Confirm the applicable cutoff timezone, live account entitlement, endpoint, API-key scope, request and task semantics, output quality, token pricing, quota, data location, retention and production support before cutover; the retirement notice itself does not specify the deadline timezone or contractual service equivalence.
Vendor source date
2026-09-02 12:00 (official announcement; service changes began 2026-09-03 00:00; documentation accessed 2026-09-05)
Primary-source verifiedVersion-pinned multimodal API model and regional availability
Alibaba Cloud releases the version-pinned qwen3.8-max-0902 snapshot across regional deployments
Alibaba Cloud Model Studio's official lifecycle catalog lists qwen3.8-max-0902, also exposed as qwen3.8-max-2026-09-02, as an upgraded Qwen3.8-Max snapshot dated September 2 for Beijing and international or global deployment scopes. The model page documents image, text and video input, text output, thinking and non-thinking modes, function calling, structured output, web search, context caching, batch inference and a one-million-token context window. The fixed snapshot ID is available alongside the rolling qwen3.8-max ID.
Buyer impact
Teams that need reproducible evaluations can test and pin the dated snapshot instead of depending only on the rolling qwen3.8-max alias. Pinning the ID does not guarantee identical regional service conditions: Alibaba's current price table lists different unit prices for Singapore than for Beijing, Virginia, Frankfurt and Tokyo, and its documentation says free quota is limited to Beijing. Before migrating, compare accepted-output quality, tool use, vision and long-video handling, latency, throughput, batch and cache behaviour; then confirm the live regional endpoint, API-key scope, account eligibility, currency, taxes, quota, data location, snapshot lifecycle and production support. Published benchmark claims are not a substitute for workload-specific testing.
Vendor source date
2026-09-02 (official snapshot and lifecycle entry; documentation accessed 2026-09-04)
Primary-source verifiedGenerated-asset storage pricing and documentation alignment
Alibaba Cloud starts billing for Model Studio asset storage above the 5 GB free quota
Alibaba Cloud Model Studio's current Chinese asset-center documentation says platform storage became generally available and entered commercial billing at 00:00 Beijing time on September 3. Each account receives 5 GB of free platform storage; usage above that allowance is billed at RMB 0.15 per GB per month and settled with Model Studio model usage. Generated images and videos are saved to platform storage by default, including assets from supported Qwen Image and Wan models. The corresponding English page still describes the asset center as invite-only beta and platform storage as free.
Buyer impact
Teams using Model Studio's generated-asset workflow should review current storage usage, billing and retention settings rather than relying on the English page's beta wording. Alibaba documents an OSS transfer option and says transferred assets stop incurring Model Studio storage charges only when the platform copy is released; items in the recycle bin still occupy billable capacity during their 30-day retention period. Before changing production workflows, confirm the live console price, account and workspace scope, region, currency, taxes, RAM permissions, OSS charges, data location and whether API-generated assets are being saved to platform storage. The documentation mismatch is a reason to preserve dated account evidence and verify the applicable terms in the intended environment.
Vendor source date
2026-09-03 (commercial billing effective at 00:00 Beijing time; official documentation accessed 2026-09-03)
Primary-source verifiedLegacy visual API retirement and platform migration
Tencent stops new purchases for six legacy Hunyuan visual APIs ahead of September 30 shutdown
Tencent Cloud's official retirement notice says new purchases stopped on September 1 for six legacy Hunyuan specialist interfaces: TextToImage, TextToImageLite, HunyuanTextToVideo, HunyuanImageToVideo, HunyuanImageChat and HunyuanTo3D. Tencent schedules full service termination for September 30 and points customers to TokenHub models HY-Image-3.0, HY-Video-1.5 and HY-3D-3.1. Tencent's migration guide documents a different TokenHub API key and base URL, while the current Hy video guide exposes hy-video-v1.5 through an asynchronous TokenHub generation endpoint.
Buyer impact
New procurement through the affected legacy interfaces is no longer available. Existing users should identify every affected endpoint and downstream task-status or download dependency, then complete a controlled TokenHub replacement test before September 30. Do not treat the recommended models as drop-in replacements: verify request fields, task lifecycle, output controls, accepted-output quality, latency, account entitlement, current pricing, quota, data location and production support. The notice does not state a cutoff timezone, contractual migration equivalence or service-level continuity, so preserve dated account evidence and allow time for a staged cutover.
Vendor source date
2026-08-28 announcement; new purchases stopped 2026-09-01; full service shutdown scheduled for 2026-09-30
Primary-source verifiedRegional API model availability and documentation gap
Alibaba Cloud lists qwen-flash-character for global deployment from Virginia
Alibaba Cloud Model Studio's official lifecycle catalog records qwen-flash-character as a Global-scope role-playing model in US (Virginia) on August 30. The dedicated model page, updated August 31, describes a text-only multilingual role-playing model with an 8,192-token context window and context caching, while marking function calling and structured outputs as unsupported. That page currently documents capability tables for Beijing and Singapore, and the current pricing page lists qwen-flash-character prices for Beijing and Singapore but not in its Virginia role-playing section.
Buyer impact
Teams considering a US-hosted role-playing model can add qwen-flash-character to a controlled availability check, but should not treat the lifecycle row alone as proof that their Virginia account can invoke and bill the model. The regional catalog, dedicated model page and pricing table are not yet fully aligned. Confirm the live model list, endpoint, API-key scope, account entitlement, currency, price, quota, web-search behaviour, data location and production support in the intended Virginia workspace before procurement or integration. The model is not a general agent replacement because the official capability table does not support function calling or structured outputs.
Vendor source date
2026-08-30 (US Virginia catalog entry; model page updated 2026-08-31)
Primary-source verifiedAPI model and capability retirement
Kling schedules legacy API models and 119 video effects for retirement on September 15
Kling AI's official API documentation says selected legacy services will be unavailable after September 15, 2026. The affected list includes Kling Image 1.0, 1.5, 2.0 and 2.0 New; Kling video 1.0, 1.5, 1.6, 2.0 Master, 2.1 and 2.1 Master; the Virtual Try-On API; and 119 Video Effects templates. Kling states that content generated before the retirement date will remain available.
Buyer impact
Teams using any listed model, Virtual Try-On or a Video Effects template should inventory production calls now, identify every affected model or effect identifier, and run a controlled replacement test before the deadline. Do not assume a newer Kling model is behaviourally or commercially equivalent: compare request fields, supported modes, output quality, latency, unit deductions and account access in the intended region. Kling recommends Image 3.0 or 3.0 Omni for Virtual Try-On use cases, but it does not state that these are drop-in replacements. The notice does not specify a cutoff timezone, migration credits, price protection or service-level terms; confirm those details with the live account and current documentation.
Vendor source date
Undated official notice accessed 2026-08-31; retirement effective 2026-09-15
Primary-source verifiedOpen-weight language model and deployment tooling
Tencent open-sources Hy4 preview with 1M context and an FP8 deployment path
Tencent Hunyuan has released Hy4 preview and an FP8 variant under the Apache 2.0 license. The official model card documents a 770-billion-parameter mixture-of-experts backbone with 49 billion parameters activated per token, a one-million-token context window, and an additional native multi-token-prediction layer for speculative decoding. Tencent provides vLLM and SGLang deployment recipes plus an OpenAI-compatible local serving example.
Buyer impact
Teams evaluating self-hosted Chinese models can now inspect, benchmark and fine-tune an official Hy4 preview release instead of relying only on a hosted endpoint. Treat it as a large infrastructure project: the official FP8 examples use tensor parallelism across eight GPUs, and Tencent explicitly describes the release as an early version with over-reasoning and over-verification issues. The open-weight release does not establish managed-service pricing, regional availability, service levels, data location or production support; verify those separately before procurement or migration.
Primary-source verifiedMultimodal API model and regional availability
Alibaba Cloud Model Studio adds Qwen3.8-Flash to Beijing and international deployment
Alibaba Cloud Model Studio's official model lifecycle catalog records qwen3.8-flash on August 26 for China (Beijing) and for a deployment scope labelled International. The dedicated model page documents the fixed qwen3.8-flash invocation ID, image, text and video input, thinking and non-thinking modes, OpenAI- and Anthropic-compatible interfaces, and a one-million-token context window. The current pricing documentation lists separate Beijing and Singapore prices, while the two deployments have different batch support and published throughput limits.
Buyer impact
Teams using Alibaba Cloud can now add the fixed qwen3.8-flash model ID to a controlled shortlist in Beijing or an international Model Studio setup. Protocol compatibility may reduce initial integration work, but it does not establish drop-in behavioural parity with another provider. Retest vision and long-video inputs, tool use, thinking controls, long-context quality, latency, throughput and output limits in the intended workload. The International label should not be read as universal availability: confirm the live region, endpoint, API-key scope, account eligibility, pricing, free quota, caching and batch rules, data location and production support before procurement or migration.
Vendor source date
2026-08-26 (official model catalog entry; documentation accessed 2026-08-29)
Tencent Cloud's August 24 TokenHub product update adds the DeepSeek-V4-Flash 0731 and DeepSeek-V4-Pro 0813 general-availability models. The current language-model overview exposes version-pinned IDs for both releases alongside rolling aliases, while the DeepSeek integration guide documents access through TokenHub's OpenAI-compatible chat endpoint. TokenHub's pricing page lists separate peak and off-peak RMB rates for the two vendor-direct models.
Buyer impact
Teams that prefer a Tencent Cloud account can now test or pin these specific DeepSeek V4 releases without relying only on a floating model alias. Confirm the exact model ID before migration because TokenHub also lists older or rolling DeepSeek V4 routes. The official guide says the vendor-direct service is supplied by DeepSeek, carries no TokenHub SLA, and is not covered by TokenHub's service agreement; users must accept DeepSeek's terms. Pricing is time-dependent, and account access, region, quotas, taxes, data handling and production support should be verified in the intended Tencent Cloud environment before purchase or deployment.
DeepSeek's current pricing page defines peak hours only from Monday through Friday and states that off-peak rates are half of peak rates. This makes Saturday and Sunday in Beijing time entirely off-peak under the schedule reported effective from August 23, 2026. The listed weekday peak windows remain 01:00-04:00 and 06:00-10:00 UTC; all other hours are off-peak.
Buyer impact
Buyers using DeepSeek's direct API can move flexible batch and evaluation workloads to the Beijing-time weekend without encountering a peak-price window. Do not translate the rule into a local weekend without checking the timezone: Beijing is UTC+8, so the recurring weekend window runs from Friday 16:00 UTC through Sunday 16:00 UTC. This is a scheduling change, not a new base-price cut. Verify the live pricing page, model, account access, taxes, credits and third-party host pricing before committing spend.
Vendor source date
2026-08-23 (effective date; official pricing accessed 2026-08-24)
DeepSeek adds an experimental V4 Flash vision model to its API
DeepSeek's current API documentation lists deepseek-v4-flash-vision-exp as an experimental model that accepts images with text. The official Vision guide documents JPEG, PNG, GIF and WebP inputs through base64, external URLs or the Files API, with support across Chat Completions, Responses and the Anthropic-compatible endpoint. DeepSeek's pricing page lists the model at the same peak and off-peak token rates as V4 Flash and states that image inputs are converted to billable tokens.
Buyer impact
Teams needing screenshot, chart or image analysis can now run a controlled direct-API evaluation without adding a separate vision provider. Treat the model as experimental rather than production-stable: the official pages do not publish a general-availability date, service-level commitment or region-by-region availability. Images are resized for tokenization and capped at 384 input tokens each, while file size, request size and image-count limits also apply. Verify current pricing, account access, provider terms, data handling and workload-specific quality before procurement or production use.
Vendor source date
2026-08-21 (first reliable report; official documentation accessed 2026-08-23)
Kuaishou's August 19 results announcement states that Kling AI officially rolled out native 4K video output in the Kling 3.0 series, released Kling 3.0 Turbo, and launched Kling MCP and Kling CLI for agent-orchestrated batch creation. Kling's current product surface also lists Video 3.0 Turbo as a faster generation option and presents 4K as an available mode.
Buyer impact
Teams evaluating Kling should retest output quality, generation speed, workflow orchestration and total cost before choosing a production path. The official results announcement and current product surface do not specify API model IDs, endpoint support, plan eligibility, pricing, quotas, licensing terms or regional availability. Native 4K shown in the creator product should not be assumed to be available through every self-service API account; verify the intended account and region before procurement or integration.
Primary-source verifiedMedia processing API controls
Tencent Cloud Media Processing adds image-quality assessment and generation tuning fields
Tencent Cloud's official Media Processing API update history shows that the August 20 schema release added ImageQualityConfig to ImageTaskInput. The corresponding data-structure documentation lists Brightness, Contrast, Sharpness and IQA as selectable assessment dimensions. An August 18 schema release also added AdditionalParameters to CreateImageConfig; Tencent's example passes a JSON string such as {"topP":0.6} to WAND-create image-generation tasks.
Buyer impact
Teams evaluating Tencent Cloud Media Processing can now design explicit image-quality gates and pass model-specific generation controls through the documented schema instead of relying only on manual review. Treat this as an API-surface update, not a new Hunyuan model release. The official pages do not state pricing, quotas, regional availability or whether every account has production access, so buyers should verify those items in the intended Tencent Cloud region before deployment.
Vendor source date
2026-08-20 (latest API schema update; related field added 2026-08-18)
DeepSeek sets peak and off-peak V4 API prices effective August 16
DeepSeek's official pricing page now states that peak and off-peak billing for deepseek-v4-flash and deepseek-v4-pro will take effect at 16:00 UTC on August 16, 2026. Peak hours are 01:00-04:00 and 06:00-10:00 UTC. V4 Flash is listed at $0.007/$0.014 for cached input, $0.22/$0.44 for uncached input, and $0.66/$1.32 for output per 1M tokens at off-peak/peak rates. V4 Pro is listed at $0.022/$0.044, $0.66/$1.32, and $1.98/$3.96 respectively.
Buyer impact
Direct DeepSeek API buyers should recalculate production budgets before the cutover and separate latency-sensitive traffic from schedulable batch work. Off-peak scheduling may reduce unit cost, but the new rates are still higher than the current flat prices. These figures apply to DeepSeek's own API pricing page; third-party hosts, regional access, taxes, credits, contractual terms, and live availability may differ and should be checked at purchase time.
DeepSeek lists V4 Pro 0813 with a 1M-token context window and a pricing increase warning
DeepSeek's official API pricing page now lists deepseek-v4-pro as version DeepSeek-V4-Pro-0813. It supports thinking and non-thinking modes, a 1M-token context window, up to 384K output, OpenAI-compatible and Anthropic-compatible endpoints, and a concurrency limit of 500. The page currently lists $0.003625 per 1M cached input tokens, $0.435 per 1M uncached input tokens, and $0.87 per 1M output tokens, while warning that overall DeepSeek API pricing is expected to rise significantly subject to an official notice.
Buyer impact
API buyers can begin controlled V4 Pro evaluations without changing the DeepSeek base URL, but should treat the current unit prices as provisional. Benchmark quality, latency, output-token usage, and concurrency against V4 Flash, and avoid hard-coding production budgets until DeepSeek publishes its final pricing notice.
Qwen Image 3.0 Pro is the latest documented Model Studio image release
Alibaba Cloud Model Studio lists Qwen Image 3.0 Pro as an invite-only preview for the Chinese mainland. The official entry highlights long-text input, dense image-in-image layouts, multilingual text rendering, and high-fidelity interface simulation.
Buyer impact
English-market buyers should treat this as a capability signal, not a procurement-ready international API. Confirm regional availability, access eligibility, pricing, output rights, and production support directly with the provider before shortlisting it.
We use optional analytics only with your permission. Essential site functions work without analytics. You can change this choice from the privacy page.