DeepSeek review
Updated 2026-08-27Product & capabilities
DeepSeek is a general-purpose AI assistant on the web and mobile, an API platform for developers, and an open-weight model family that can be self-hosted. The current flagship is DeepSeek V4 Pro 0813. DeepSeek now states that V4 Pro is available through the API, app and web product, so the older assumption that the consumer assistant and the current API model must be reviewed as separate generations is no longer valid.12
The official consumer assistant remains free. It offers instant and expert-style reasoning modes, web search and file handling. Developers pay by token for the hosted API, with peak and off-peak pricing. Teams can also download the V4 weights under the published license and run them on infrastructure they control, although the scale of the Pro model makes that a serious engineering project rather than a desktop-friendly privacy switch.34
This is a desk review. AiToolMap did not conduct a controlled hands-on DeepSeek test in this cycle. Product behavior and quality claims are attributed to current first-party documentation, the fixed 50-source panel and additional independent evidence.
Artificial Analysis gives V4 Pro 0813 a score of 53 on its Intelligence Index. It also reports a one-million-token context window, 68.3 output tokens per second and a 1.48-second time to first token for the tested provider configuration.5
Those numbers make V4 Pro a credible frontier-adjacent model, particularly at its price. They do not prove that the web and mobile assistant will deliver the same latency or quality in every region, because product performance also depends on search, routing, moderation, load and interface behavior.
LM Arena supplies a second signal, but a much less stable one. Its current Expert leaderboard shows V4 Pro 0813 around a score of 1479 with a wide uncertainty band and only 353 votes.6 AiToolMap treats that as preliminary preference evidence, not as a precise product rank.
Epoch places the V4 Pro preview lineage at ECI 149, rank 43 of 229 models.7 The checkpoint predates the 0813 general-availability release, so it is retained as lineage context rather than silently merged with the current flagship.
ARC Prize reports strong ARC-AGI results for V4 Flash 0731: 89.0% on ARC-AGI-1 and 61.4% on ARC-AGI-2 at very low task cost.8 That supports the broader V4 family’s reasoning efficiency, but Flash is not Pro and ARC-AGI is not a finished-assistant review.
The combined conclusion is careful but positive: current DeepSeek models are highly capable, especially for reasoning and coding relative to cost. The evidence does not justify turning one leaderboard into an overall product score.
TechRadar’s current DeepSeek review is the strongest accessible hands-on source in the panel. Its reviewer used the free web app for coding, document summarization, mathematics, multi-turn questions and file analysis, and also exercised the API. The verdict was favorable on capability, price and open weights, with serious reservations about data residency, security history and censorship.11
The official app-store populations show that the product is not a niche developer interface. Google Play lists more than 50 million downloads and an English-store rating around 4.1/5 from roughly 316,000 reviews. The US App Store is around 4.0/5 from approximately 11,000 ratings.1213
App-store feedback is mixed in recognizable ways. Positive reviews value free access, speed, coding help and strong direct answers. Repeated complaints include losing conversational context, hallucinations, weak instruction-following, unexpected switching into Chinese, censorship on politically sensitive topics, service instability and missing conveniences such as chat search or richer image features.
These samples measure the mobile experience, not only model reasoning. They are therefore useful for usability and reliability, but should not be combined mechanically with benchmark scores.
G2 shows roughly 4.7/5 from 14 reviews. Reviewers commonly praise speed, ease of use, value, coding and research. Negatives include privacy concerns, accuracy on complex context, limited integrations, censorship and support. The sample is small, several reviews were collected through G2 incentives, and the reviews sometimes blur the app, API and underlying model.14
TrustRadius shows 8.1/10 from nine ratings.15 Its current product description still references the older 67-billion-parameter DeepSeek LLM, which creates a surface and recency mismatch. The rating is preserved as a professional-user signal but receives limited weight.
Trustpilot is sharply more negative at roughly 2.5/5 from 151 reviews. Positive comments mention coding, speed and free access; negative comments emphasize hallucinations, censorship, instruction-following, language switching, sign-up and support problems.16 Under AiToolMap’s standing methodology, Trustpilot is qualitative only and excluded from the direct numeric equation.
PeerSpot has only one old pre-V4 review and is excluded. Gartner has a current listing but no reviews. Capterra has an exact-name pricing page without a rating population, and its asserted subscription tiers conflict with the first-party usage-based schedule; those claims are not used.
Artificial Analysis labels V4 Pro open weights, and DeepSeek publishes downloadable model artifacts under a permissive license.51 This gives technical teams options that a fully closed assistant does not: inspect deployment behavior, choose a serving layer, keep data inside their own environment and adapt the surrounding agent stack.
“Open weights” should not be inflated into “fully open development.” Training data, data governance and the complete training process are not reproduced simply because weights are downloadable.
The infrastructure requirement is also substantial. Simon Willison reports roughly 893 GB of weights for V4 Pro and around 167 GB for V4 Flash.4 The smaller Flash model is the more realistic self-hosting target for many organizations. Pro-scale local deployment requires serious accelerators, memory, serving expertise, monitoring and security controls.
Censorship appears consistently across TechRadar, G2, app-store reviews and Trustpilot. The practical issue is not limited to political discussion. A system that refuses, reframes or changes language unpredictably can be unreliable for research, journalism, policy analysis and international business.
This limitation should be separated from hallucination. Censorship is a policy boundary; hallucination is an accuracy failure. Both can reduce usefulness, but they require different mitigations and should not be compressed into one generic “safety” label.
DeepSeek is best for cost-sensitive developers, technical users, coding and reasoning workloads, and consumers who want a powerful free assistant. It is especially attractive when open weights, self-hosting options or very low API cost are important.
It is a weaker default for sensitive regulated work, organizations that need mature administrative controls and contractual assurances, and users who depend on politically unrestricted research or a deep productivity-suite ecosystem.
Pricing & access
Consumer chat is currently free. There is no confirmed paid consumer subscription tier in the official material reviewed for this cycle.1
The API uses a time-dependent schedule. For V4 Pro, current peak pricing is $1.32 per million input tokens and $3.96 per million output tokens; off-peak pricing is $0.66 input and $1.98 output. V4 Flash is $0.44 input and $1.32 output at peak, and $0.22 input and $0.66 output off peak.31017
The August price increase made several otherwise current July reviews stale on cost. Even after the increase, DeepSeek remains inexpensive relative to many leading hosted models, but buyers should not assume that launch promotions or old price tables will persist. Interactive workloads also cannot always be shifted into cheaper off-peak windows.
For high-volume teams, the economic advantage is real. For regulated or sensitive work, lower token cost should be evaluated alongside residency, policy, vendor support and the operational cost of self-hosting.
Evidence & trust
This review was rebuilt under methodology v1.8 after a substantive audit of all 50 members of ai-review-panel-2026-08-v4. Twenty-six sources produced current material, ten produced older or checkpoint-mismatched material, five produced no relevant exact-product result, and nine were inaccessible in the audit environment. Stale benchmarks and blocked or paywalled pages were not promoted from snippets into evidence.
The strongest current evidence is a combination of exact-checkpoint model benchmarks, a recent hands-on product review, large app-store populations, professional-review samples, current launch and pricing reporting, and the company’s present privacy, status and API documentation. The evidence is still uneven. Model quality is documented much more deeply than the reliability of the finished assistant across long tool-using workflows.
No public AiToolMap score is assigned yet. Rating inputs have been captured before corpus-wide normalization, missing sources are not treated as zero, G2 and Capterra remain one independence family, and Trustpilot is qualitative only.
DeepSeek positions V4 Pro as materially improved for agent capabilities and has released an open Harness built around replaceable plugins for models, tools, sessions, sandboxes, filesystems and orchestration.19 This is attractive for developers who want to inspect or replace more of the stack than a closed assistant normally permits.
Real-task evidence shows why the harness matters. VentureBeat reports that V4 Flash completed 53.8% of 240 runs across 30 difficult, multi-step tasks when tested through eight agent harnesses. Results varied materially with orchestration, tools, caching, retries and provider configuration.10
That test covers Flash rather than V4 Pro, so it cannot be used as a direct Pro failure rate. It is still an important warning against assuming that a strong model score automatically produces a production-ready agent. For enterprise workflows, the serving provider, harness, permissions, state management and recovery behavior are part of the product.
DeepSeek’s current privacy policy says it collects account information, prompts and other user input, including uploads, voice or photos where used, along with device, network, log, location and payment-related data. It states that directly collected information is processed and stored in the People’s Republic of China and may be shared with service providers and corporate-group companies, including for model training and optimization.18
The policy also describes retention while an account exists and for legal or business needs, and provides rights such as access, deletion, restriction and portability. It says the service is not designed for sensitive personal data.
This does not mean every DeepSeek deployment has identical risk. A self-hosted open-weight model can have a very different data path from the hosted consumer app. A third-party model host can have a different policy again. Buyers must evaluate the exact route: official app, official API, third-party API or self-hosted weights.
For confidential client data, regulated records, trade secrets or personal data, the default consumer app should not be treated as an enterprise privacy boundary without a separate legal, security and residency review.
DeepSeek’s public status history for the current review window shows reported uptime around 99.76% to 100% across key services, including the Pro API, Flash API, chat, search and file/vision features.19
A status page is a first-party operational record, not an independent reliability benchmark. It is still more current and specific than recycling outage reports from the R1 launch period. User reviews continue to report intermittent slowness and failed responses, so the evidence supports “generally available and broadly reliable” rather than “problem-free.”
Strengths & weaknesses
Strengths
The web and mobile assistant offers unusually strong reasoning, search and file work without a consumer subscription.
Independent model evidence, hands-on testing and user reviews consistently support coding, mathematics and structured analytical work.
Even after a significant price increase, V4 Pro and especially V4 Flash remain compelling for workloads where cost per useful result matters.
Downloadable weights and an open Harness create more control over serving and orchestration than closed assistants offer.
The one-million-token model context is attractive for large documents and codebases, subject to provider and product limits.
Weaknesses
Official consumer data is stored in China and the policy permits broad collection and sharing. Organizations need route-specific review before using confidential data.
Multiple independent user surfaces report refusals, politically constrained answers and unexpected Chinese output.
DeepSeek is less mature than the largest generalist ecosystems in native productivity integrations, multimodal creation, memory, voice and enterprise administration.
Strong leaderboards do not eliminate orchestration failures. Current real-task evidence for V4 family agents is mixed and highly harness-dependent.
Professional review samples are small, enterprise support evidence is limited and there is no broad current Gartner/PeerSpot population.
Rapid changes make cached comparison tables unreliable and complicate long-term budgeting.
Sources & references
- Official sourceDeepSeek — official product and V4 Pro announcementOFFICIAL
- ReutersReuters — DeepSeek launches V4 ProNEWS2026-08-13
- Official sourceDeepSeek API — Models and PricingOFFICIAL
- Simon WillisonSimon Willison — DeepSeek V4 Pro 0813EXPERT ANALYSIS2026-08-12
- Artificial AnalysisArtificial Analysis — DeepSeek V4 Pro 0813BENCHMARK2026-08-13
- LM ArenaLM Arena — Expert LeaderboardBENCHMARK
- Epoch AIEpoch AI — DeepSeek V4 ProBENCHMARK2026-04-24
- ARC PrizeARC Prize — DeepSeek V4 Flash 0731BENCHMARK2026-07-31
- The RegisterThe Register — DeepSeek HarnessEXPERT ANALYSIS2026-08-14
- VentureBeatVentureBeat — V4 Flash real agent tasksEXPERT ANALYSIS2026-08-16
- TechRadarTechRadar — DeepSeek AI reviewEDITORIAL REVIEW2026-07-13
- Google PlayGoogle Play — DeepSeek AI AssistantUSER REVIEWS
- Apple App StoreApple App Store US — DeepSeek AI AssistantUSER REVIEWS
- G2G2 — DeepSeek ReviewsUSER REVIEWS
- TrustRadiusTrustRadius — DeepSeekUSER REVIEWS
- TrustpilotTrustpilot — DeepSeekUSER REVIEWS
- EngadgetEngadget — DeepSeek’s V4 price changeNEWS2026-08-14
- Official sourceDeepSeek — Privacy PolicyOFFICIAL
- Official sourceDeepSeek StatusOFFICIAL