Google Gemini 3.5 Flash
A fast Gemini 3 model: frontier-level reasoning at Flash-level latency and cost, tuned for agentic workflows and iterative coding.
From Google, Gemini 3.5 Flash is a cloud text model. It is multimodal: alongside a text prompt it accepts images, video, and audio, then replies with generated text. Its context window handles up to 1.05M input tokens and up to 66K output tokens, and it supports adjustable reasoning effort for harder problems. Generation can be tuned with system instructions, temperature, and top-p sampling. It runs through Runware and Replicate using your own API key, from $1.50 per million input tokens.
- Released
- 4 months ago (May 2026)
- Pricing
- $1.50 / 1M in · $9.00 / 1M out
- Inputs
- Text, Images, Video, Audio
- Context window
- 1.05M in · 66K out
- Reasoning
- Adjustable effort
- Controls
- System prompt, Temperature, Top-p
Examples
Generated with Gemini 3.5 Flash via Replicate. The same three prompts run against every text model in the catalog, shown verbatim — so the only thing that changes between two models’ answers is the model.
Explain what an API is to someone who has never written code. Do not use a restaurant or waiter analogy. Plain prose only — no headings, bullet points, or markdown formatting. Maximum 50 words.
An API is a digital bridge that lets different software programs talk to each other. When you use a weather app, it uses an API to request data from a weather service and show it on your screen. It is a translator that helps separate apps share information.
A two-person startup is deciding whether to build their own user authentication or use a hosted provider. Give them a clear recommendation, then the single strongest argument against your own recommendation. Plain prose only — no headings, bullet points, or markdown formatting. Maximum 150 words.
As a two-person startup, you should use a hosted authentication provider. At your stage, speed to market and focus on your core product are everything. Building a secure, compliant, and feature-rich authentication system takes weeks of engineering time and requires ongoing maintenance that distracts from your primary value proposition. Hosted services let you launch securely in hours. However, the strongest argument against this recommendation is the risk of long-term vendor lock-in and unpredictable scaling costs. Once your user base grows, migrating away from a proprietary provider is a technical nightmare that risks disrupting your users. You may find yourself captive to their pricing model, paying exorbitant fees for a utility that would have been virtually free to run on your own servers once established.
Write a briefing for a non-technical executive explaining why their company's spending on AI is hard to predict, and what they can do about it. Cover at least three distinct causes. Plain prose only — no headings, bullet points, or markdown formatting. Write exactly 205 words.
Managing artificial intelligence expenses are uniquely challenging because this technology does not follow traditional software cost structures. First, computing power demands are highly volatile; as your employees query these systems, backend usage costs can spike exponentially based on the complexity of the data processed rather than just the number of active users. Second, the commercial market is in constant flux, with major vendors frequently changing their consumption structures, token rates, and subscription models overnight without warning. Third, development is fundamentally experimental, meaning software engineers must often spend unpredictable amounts of money testing models that ultimately fail before finding a viable solution. Fortunately, you can regain control over these fluctuating corporate budgets with a few very effective adjustments. You should immediately implement strict application programming interface spending caps and automated alerts to prevent runaway costs from daily operations. Additionally, establish a disciplined framework that treats initial pilots as highly budgeted research experiments rather than standard, predictable technology rollouts. Finally, prioritize smaller, specialized models tailored for your specific business tasks rather than relying on massive, expensive general purpose systems. By treating artificial intelligence as a dynamic, metered utility rather than a fixed capital asset, your business can easily protect its quarterly bottom line and overall fiscal health.
Google and Google DeepMind build the Gemini family of multimodal models, the Imagen and Nano Banana image models, the Lyria music models, and the Veo video models.
deepmind.google ↗