Skip to content
modsignal

Perplexity

perplexity.ai · AI · watching since Aug 17, 2026 · RSS

Follow
https://docs.perplexity.ai/changelogChecked every 168 h · last checked 1 d ago · watching since Aug 17, 2026
Watch it yourself →

Current state

As of Sep 6, 2026

Perplexity's API changelog tracks frequent additions of new third-party and in-house models (e.g., GLM 5.3, Grok 4.6, NVIDIA Nemotron variants) to the Agent API and Router API, alongside pricing updates, preset changes, and infrastructure additions like the Router API and remote MCP Server. Several older Sonar and Gemini models have been deprecated and removed in favor of newer replacements such as Sonar Pro and Agent API presets.

Deprecations

  • sonar-reasoning model

    sonar-reasoning-pro

    Sunset 2025-12-15
  • google/gemini-2.5-flash

    Sunset 2026-03-20
  • google/gemini-2.5-pro

    Sunset 2026-04-01
  • google/gemini-3-pro-preview

    Sunset 2026-04-01
  • google/gemini-3.1-flash-lite-preview

    google/gemini-3.1-flash-lite

    No date
  • llama-3.1-sonar-small-128k-online, llama-3.1-sonar-large-128k-online, llama-3.1-sonar-huge-128k-online

    Sonar or Sonar Pro

    Sunset 2025-02-22
  • llama-3-sonar-small-32k-online, llama-3-sonar-large-32k-online, llama-3-sonar-small-32k-chat, llama-3-sonar-large-32k-chat, llama-3-8b-instruct, llama-3-70b-instruct, mistral-7b-instruct, mixtral-8x7b-instruct

    Llama-3.1 family models

    Sunset 2024-08-12
  • codellama-70b-instruct, mistral-7b-instruct, mixtral-8x22b-instruct, pplx-7b-chat, pplx-7b-online

    Llama-3 family models

    Sunset 2024-05-14
  • Sonar Chat Completions

    Agent API

    No date

Latest entries

    • GLM 5.3 Flash added to Agent API and Router APIaddition
    • GLM 5.3 added to Agent API and Router APIaddition
    • Prompt caching enabled for Agent API presetsaddition
    • Fast preset updated to use gpt-5.6-luna with priority processingchange
    • Gemini 3.7 flash pricing increasedchange
    • Grok 4.6 added to Agent APIaddition
    • NVIDIA Nemotron 3 ultra added to Agent API and Router APIaddition
    • NVIDIA Nemotron 3.5 lightning added to Agent API and Router APIaddition
    • DeepSeek v4 flash 0731 added to Agent API and Router APIaddition
    • GPT-5.6 price cuts and Sol fast mode introducedchange
    • Router API launched for unified access to open-weight modelsaddition
    • Remote MCP server hosted by Perplexity now availableaddition

Changes

  1. ChangeAPI changelog

    Gemini 3.7 Flash pricing increased on August 27, 2026; GLM 5.3 Flash model added with new pricing; navigation menu reordered.

    Gemini 3.7 FlashThe Agent API and Router API now support google/gemini-3.7-flash at launch pricing of $0.375 per million input tokens, $0.0375 per million cached-input tokens, and $1.875 per million output tokens.
    Gemini 3.7 FlashPricing for google/gemini-3.7-flash increased on August 27, 2026 to $0.75 per million input tokens, $0.075 per million cached-input tokens, and $3.75 per million output and reasoning tokens. September 2026 Agent APIRouterModels GLM 5.3 FlashThe Agent API and Router API now support perplexity/glm-5.3-flash at $0.15 per million uncached-input tokens, $0.03 per million cached-input tokens, and $0.50 per million output tokens.
    Confidence95%
  2. ChangeAPI changelog

    New model GLM 5.3 added to Agent API and Router API with pricing of $1.40/$0.26/$4.40 per million tokens for uncached-input/cached-input/output respectively.

    August 2026 Agent APIPresets
    August 2026 Agent APIRouterModels GLM 5.3The Agent API and Router API now support perplexity/glm-5.3 at $1.40 per million uncached-input tokens, $0.26 per million cached-input tokens, and $4.40 per million output tokens. See the Agent API Models reference or the Router model catalog. August 2026 Agent APIPresets
    Confidence95%
  3. ChangeAPI changelog

    Perplexity added prompt caching for presets (cost reduction feature), reordered navigation menu, and clarified Agent API fast preset pricing impact (2× token costs for priority processing).

    ModelsAgent APISearchDeprecationFeaturesRouterFiltersPresetsPricingMCPMultimodalSecuritySDKToolsFinanceIntegrationsAWSBillingEmbeddingsDocsPro SearchPlaygroundFile AttachmentsAPI KeysCost TrackingUsageFinancialAcademicReasoningAsyncBreaking ChangeOrganizationPortalImagesAccessStructured OutputsCitationsRate Limits August 2026 Agent APIPresets Fast preset updated
    ModelsAgent APISearchDeprecationFeaturesPresetsRouterFiltersPricingMCPMultimodalSecuritySDKToolsFinanceIntegrationsAWSBillingEmbeddingsDocsPro SearchPlaygroundFile AttachmentsAPI KeysCost TrackingUsageFinancialAcademicReasoningAsyncBreaking ChangeOrganizationPortalImagesAccessStructured OutputsCitationsRate Limits August 2026 Agent APIPresets Prompt caching for presetsAgent API presets now use stable prompt cache keys automatically, allowing independent requests with the same preset to reuse the shared prompt prefix (system prompt and tool definitions). No request changes are required, and an explicit prompt_cache_key still overrides the preset default. This can reduce costs by about 5% for applications that use presets frequently, depending on cache utilization. The current preset values include each key for frozen configurations. August 2026 Agent APIPresets Fast preset updated
    Confidence85%
  4. ChangeAPI changelog

    New Aug 2026 entry: Agent API 'fast' preset changed to openai/gpt-5.6-luna with minimal reasoning effort and priority (2x price) processing—frozen configs must update model, reasoning effort, and service_tier.

    August 2026 Agent APIRouterModels Gemini 3.7 Flash...
    August 2026 Agent APIPresets Fast preset updatedThe Agent API fast preset now uses openai/gpt-5.6-luna with minimal reasoning effort and priority processing. Dynamic fast preset requests pick up the change automatically. If you use a frozen configuration, update the model and reasoning effort and set service_tier to priority. Priority processing uses 2× the model's standard token prices.
    Confidence97%

This is the Product and API pack running on a real vendor.

Watch Perplexity — or anything else — your way.

Your own prompt, cadence and channels. One alert with the before/after proof when your sentence comes true.

Get started