Model Risk & Observations

Risk observations

Current risks and watch items

Data as of 2026-08-10
Record state
Risk type
Next action
Source
Confidence
Matched:
Qwen3.8 Max Preview compatibility aliasAlibaba / QwenCurrent riskAPI / serviceAPI availability

Token Plan prohibits application backends, automated scripts, and non-interactive batch workloads

The qwen3.8-max-preview ID now routes to production qwen3.8-max, while the individual Token Plan terms still restrict automated scripts, application backends, and non-interactive batch workloads.

Token Plan Individual overviewOfficial · High confidenceObserved 2026-07-19Source as of 2026-08-10
Next action: Verify before deployment

Use the production model ID for deployment decisions and verify the terms of the actual service channel; the compatibility ID is not a separate version contract.

Model detailOpen source
Kimi K3Moonshot AI / KimiCurrent riskModel behaviorRuntime behavior

Incomplete historical thinking or mid-session model switching can destabilize output

Kimi requires complete historical thinking content to be passed back and warns that switching from another model to K3 mid-session can materially reduce generation stability.

Kimi K3: Open Frontier IntelligenceOfficial · High confidenceObserved 2026-07-16Source as of 2026-07-20
Next action: Validate on real tasks

Pin the model ID before migration and validate preservation and replay of thinking history in real multi-turn tool workflows.

Model detailOpen source
Kimi K3Moonshot AI / KimiCurrent riskModel behaviorRuntime behavior

Official guidance acknowledges excessive proactiveness on ambiguous tasks

Kimi's release guidance says K3 can sometimes take more action than needed for ambiguous intent or simple questions.

Kimi K3: Open Frontier IntelligenceOfficial · High confidenceObserved 2026-07-16Source as of 2026-07-20
Next action: Validate on real tasks

Production agents should constrain tool permissions, define stop conditions, and validate authorization boundaries with real tasks.

Model detailOpen source
Muse Spark 1.1MetaCurrent riskAPI / serviceRegional availability

The Model API remains a public preview currently limited to US developers

Meta's official model page marks Muse Spark 1.1 as public preview with current regional restrictions, so global teams cannot treat the announcement as broadly procurable production service.

Muse Spark 1.1 model pageOfficial · High confidenceObserved 2026-07-09Source as of 2026-07-15
Next action: Verify before deployment

Confirm account eligibility, data residency, and production SLA before cross-region rollout.

Model detailOpen source
Grok 4.5xAICurrent riskAPI / serviceCost & caching

Official docs warn that missing conversation cache keys can cause frequent full-input charges

xAI recommends prompt_cache_key or x-grok-conv-id for long conversations; otherwise requests may land on cache-cold servers and incur full input charges.

grok-4.5 | SpaceXAI DocsOfficial · High confidenceObserved 2026-07-08Source as of 2026-07-15
Next action: Verify before deployment

Cost evaluation should include cache-hit rates and compaction strategy for long agent loops.

Model detailOpen source
Claude Fable 5AnthropicCurrent riskData / governanceData retention

Safety safeguards require 30-day retention and do not support zero-data retention

Anthropic's Fable 5 safeguard documentation requires 30-day safety retention and states that zero-data retention is unavailable, directly affecting regulated data and sensitive-code deployments.

Introducing Claude Fable 5 and Claude Mythos 5Official · High confidenceObserved 2026-07-01Source as of 2026-07-15
Next action: Verify before deployment

Confirm data classification with legal and security teams; choose another model when ZDR is mandatory.

Model detailOpen source
MiMo-V2.5-Pro-UltraSpeedXiaomiCurrent riskAPI / serviceAccess restrictions

UltraSpeed remains an approval-only closed beta without a fixed end or GA date

Xiaomi extended the UltraSpeed closed beta, but access remains limited to approved users and no fixed GA date is published.

MiMo UltraSpeed beta extension noticeOfficial · High confidenceObserved 2026-06-23Source as of 2026-07-15
Next action: Verify before deployment

Latency-sensitive designs need a standard MiMo-V2.5-Pro fallback and should not depend on beta quota.

Model detailOpen source
Claude Fable 5AnthropicCurrent riskModel behaviorSafety robustness

A red-team study still observed jailbreak success under automated attacks

A preprint reported a 6.1% worst-case attack success rate for Fable 5 under its test setup. The figure applies only to that methodology and sample, not a general real-world risk rate.

A Red-Team Study of Anthropic Fable 5 & Opus 4.8 ModelsResearch · Medium confidenceObserved 2026-06-20Source as of 2026-07-15
Next action: Validate on real tasks

High-risk deployments still require independent red teaming and output isolation.

Model detailOpen source
Qwen3.5-OCRAlibaba / QwenCurrent riskAPI / serviceRegional availability

The current first-party OCR endpoint is only listed for the China region

Qwen3.5-OCR appears in Alibaba Cloud's China catalog without an equivalent globally available Model Studio endpoint. Cross-region deployment requires separate confirmation.

Alibaba Cloud Model Studio model pricingOfficial · High confidenceObserved 2026-06-16Source as of 2026-07-15
Next action: Verify before deployment

Do not assume international availability; confirm region, data residency, and contract terms first.

Model detailOpen source
Grok Imagine ImagexAICurrent riskProduct experienceRegulatory & privacy

Canada's privacy regulator found inadequate safeguards when image generation launched

Canada's privacy regulator found that X/xAI launched related image-generation capabilities without adequate privacy safeguards and added measures later. The finding concerns the deployed product, not base API capability.

Canadian privacy regulator finding on X and xAIOfficial · High confidenceObserved 2026-06-11Source as of 2026-07-15
Next action: Verify before deployment

Workflows involving real people need consent, identity, and content-review controls.

Model detailOpen source
Kimi K2.7 CodeMoonshot AI / KimiCurrent riskAPI / serviceRuntime behavior

The coding endpoint is always-thinking, so latency and output tokens cannot be estimated like a standard chat model

Moonshot's quickstart describes K2.7 Code as an always-thinking tool-use model, which can increase end-to-end latency and output cost in long agent loops.

Kimi K2.7 Code quickstartOfficial · High confidenceObserved 2026-06-11Source as of 2026-07-15
Next action: Validate on real tasks

Measure completion time and total tokens on real repository tasks, not only single-response speed.

Model detailOpen source
Claude Fable 5AnthropicCurrent riskAPI / serviceGuardrails & refusals

Safety classifiers can return refusals with HTTP 200, requiring explicit fallback handling

Official integration guidance requires applications to detect stop_reason=refusal and retry through server-side, SDK, or manual fallback paths.

Introducing Claude Fable 5 and Claude Mythos 5Official · High confidenceObserved 2026-06-09Source as of 2026-07-15
Next action: Validate on real tasks

Ignoring the refusal branch can misclassify a successful HTTP response as completed work.

Model detailOpen source
MiniMax-M3MiniMaxCurrent riskAPI / serviceContext availability

API prices for 512K-to-1M context are double the rates at or below 512K

MiniMax's enterprise API pricing page shows that input, cache, and output rates for the 512K-to-1M input range are twice the rates at or below 512K. The current page no longer supports the earlier limited-availability claim.

MiniMax enterprise API token pricingOfficial · High confidenceObserved 2026-06-01Source as of 2026-07-15
Next action: Validate on real tasks

Budget million-context workloads against the separate tier and regression-test behavior at the 512K boundary.

Model detailOpen source
LongCat-2.0MeituanCurrent riskAPI / serviceMigration risk

Six legacy LongCat Flash API entries were retired on May 29

Meituan's change log lists six legacy Chat, Thinking, and Omni model APIs as retired on May 29 and records LongCat-2.0 launching on June 30. Downloadable weights do not imply continuity for the retired APIs.

LongCat platform change logOfficial · High confidenceObserved 2026-05-29Source as of 2026-07-15
Next action: Verify before deployment

Pin LongCat-2.0 for new integrations and test model-ID, prompt, and tool-protocol migration explicitly.

Model detailOpen source
Grok 4.20 ReasoningxAICurrent riskModel behaviorSafety robustness

The official model card reports failures on agent-harm, prompt-injection, and misuse evaluations

xAI's Grok 4.20 model card reports a 30% AgentHarm violation rate, 33% AgentDojo prompt-injection success rate, and 32% misuse violation rate under system-prompt override. These figures apply only to the documented tests.

Grok 4.20 model cardOfficial · High confidenceObserved 2026-04-07Source as of 2026-07-15
Next action: Validate on real tasks

Do not use the model directly for autonomous high-risk decisions; minimize and isolate agent tool permissions.

Model detailOpen source
Grok 4.5xAIWatchProduct experienceUsage limits

A new report says SuperGrok 4.5 hit its weekly allowance after roughly two hours of coding

In a July 14 r/grok thread, one user said a routine JavaScript coding session through Grok CLI reached the SuperGrok 4.5 weekly allowance in roughly two hours. The same reply praised speed but reported needing more prompts and questioned plan transparency. This single public thread concerns the consumer subscription/Grok CLI experience, not xAI API limits or a representative sample.

Grok weekly limitsCommunity · Low-confidence leadObserved 2026-07-14Source as of 2026-08-04
Next action: Monitor

Validate subscription allowances and prompt counts on real workflows; plan production capacity against a pinned API model and documented API limits.

Model detailOpen source
Grok 4.5xAIWatchProduct experienceQuality regression

A recent user thread reports declining accuracy and writing experience

A recent r/grok post alleges more frequent incorrect information, with replies also reporting weaker writing quality. A single thread is only a lead for validation.

So long Grok. It was fun, then it was awful.Community · Low-confidence leadObserved 2026-07-12Source as of 2026-08-04
Next action: Monitor

Run regression tests on real workloads; do not reject the model from one community thread alone.

Model detailOpen source
Claude Fable 5AnthropicWatchProduct experienceGuardrails & refusals

Security researchers reported guardrails triggering on routine code review

TechCrunch summarized reports from security researchers that some routine code-review and security-reading requests triggered Fable safeguards or fallback behavior.

Cybersecurity researchers aren't happy about the guardrails on Anthropic's FableNews · Medium confidenceObserved 2026-06-10Source as of 2026-08-04
Next action: Validate on real tasks

This is feedback from a specific professional cohort, not all workloads; security teams should run workflow-specific trials.

Model detailOpen source
Record stateCurrent riskStill affects selection, procurement, or deployment on the date shown on the card.WatchPublic information is changing or incomplete and is not yet a firm conclusion.Historical recordPreviously active, expired, or no longer sufficient for a current risk judgment; excluded from the default current list.
Next actionVerify before deploymentBefore integration or procurement, check the current endpoint, model ID, and usage conditions.Validate on real tasksUse your own workload to validate capability and stability.MonitorEvidence is still changing or incomplete; continue tracking it.
Historical records7
Qwen3.8 Max Preview compatibility aliasAlibaba / QwenHistoricalAPI / serviceVersion transparency

The preview model ID now routes to the production model

QwenCloud now confirms that qwen3.8-max-preview remains callable but automatically routes to production qwen3.8-max, with billing and usage statistics based on the production model.

Token Plan Individual overviewOfficial · High confidenceObserved 2026-07-19Source as of 2026-08-10Ended 2026-08-10
Next action: Historical record

Keep dates on historical preview evaluations; current selection and cost decisions should use qwen3.8-max rather than treating the preview ID as a separate model.

Model detailOpen source
Kimi K3Moonshot AI / KimiHistoricalAPI / serviceAccess restrictions

Full weights were published on July 27, closing the earlier availability gap

Moonshot has published the full Kimi K3 weights, model card, and Kimi K3 License; the earlier unavailability observation is no longer a current risk.

MoonshotAI/Kimi-K3 official repositoryOfficial · High confidenceObserved 2026-07-16Source as of 2026-08-04Ended 2026-07-27
Next action: Historical record

Local deployment still requires review of the Kimi K3 License; the open weights do not include Kimi.com's hosted agent orchestration, tool routing, or serving stack.

Model detailOpen source
DeepSeek V4 FlashDeepSeekHistoricalAPI / serviceAlias & lifecycle

The July 24 migration deadline has passed, so the forward-looking alert is historical

The original record described a pre-deadline migration risk and cannot remain a current negative claim after the cutoff. Current alias targets and available model IDs require fresh verification against the live API model list.

DeepSeek current API model documentationOfficial · High confidenceObserved 2026-07-13Source as of 2026-08-04Ended 2026-07-24
Next action: Historical record

Production calls should still pin explicit model IDs, record the served version, and refresh the current model list before deployment.

Model detailOpen source
Grok Imagine Video 1.5xAIHistoricalProduct experienceUsage limits

A July 10 thread reported that the 15-second generation option was unavailable

The source supports only that one user could not select the 15-second generation option at that time; it does not establish that older assets were deleted. The scope was corrected on July 15 and moved to history.

xAI reset Grok limits and removed 15s from ImagineCommunity · Low-confidence leadObserved 2026-07-10Source as of 2026-07-15Ended 2026-07-15
Next action: Historical record

Product limits and options can change; retain important assets in owned storage.

Model detailOpen source
Seedream 5.0 ProByteDance / DoubaoHistoricalProduct experienceAPI availability

Callable API availability was confirmed on July 15, closing the earlier evidence gap

The official Seedream 5.0 Pro page now provides a Get API entry linked to the corresponding BytePlus Ark model listing. First-party evidence corrects the earlier unverified-API observation, so it is historical.

Seedream 5.0 Pro official model pageOfficial · High confidenceObserved 2026-07-08Source as of 2026-07-15Ended 2026-07-15
Next action: Historical record

For procurement, still verify pricing, quota, and SLA for the account region.

Model detailOpen source
Grok 4.5xAIHistoricalProduct experienceVersion transparency

A July 2 thread questioned model-version transparency in the product UI

The thread alleged a rollback or difficulty confirming the served model, but model self-identification and one user's inference do not establish current routing. No sufficient evidence was found on July 15 to keep it as a current risk, so it is historical.

Grok model version roll-back to 4.20?Community · Low-confidence leadObserved 2026-07-02Source as of 2026-07-15Ended 2026-07-15
Next action: Historical record

For serious evaluations, continue to pin API model IDs and record response metadata.

Model detailOpen source
Claude Fable 5AnthropicHistoricalAPI / serviceAvailability

Access was suspended under export controls and restored globally on July 1

Access was suspended on June 12 and restored globally on July 1. The interruption ended and is retained as a supply-continuity record.

Redeploying Claude Fable 5Official · High confidenceObserved 2026-07-01Source as of 2026-07-15Ended 2026-07-01
Next action: Historical record

No active outage response is required; retain provider-switching and model-fallback plans.

Model detailOpen source