If you already use the OpenAI Python SDK, you already know how to call a model:
from openai import OpenAI
client = OpenAI(
api_key="sk-or-...",
base_url="https://api.openai.com/v1"
)
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": "What changed in EU AI regulation this week?"}]
)
print(response.choices[0].message.content)
Now add one thing: live web search. With most OpenAI-compatible gateways, you don’t need tools, a search key, or an agent loop. You only change the base_url.
The usual way: tools=[{type: “web_search”}]
OpenAI’s native web search requires the Responses API and an explicit tool declaration:
from openai import OpenAI
client = OpenAI(
api_key="sk-...",
base_url="https://api.openai.com/v1"
)
response = client.responses.create(
model="gpt-4o",
tools=[{"type": "web_search"}],
input="What changed in EU AI regulation this week?"
)
print(response.output_text)
Two requirements follow:
- You must use the Responses API, not the classic Chat Completions endpoint.
- You must pass
tools=[{"type": "web_search"}]on every request — there is no default.
What that costs and adds
- A separate
search:browsepermission scope - A Tavily or Bing key if you are not on OpenAI’s plan
- A
max_search_resultsvalue you tune by hand - No upgrade path to other models without another SDK import
The one-line way
Point the same OpenAI SDK at a gateway that runs search automatically:
from openai import OpenAI
# Same OpenAI SDK you already use
client = OpenAI(
api_key="your-easyrouterai-key",
base_url="https://easyrouterai.com/v1"
)
# No tools parameter, no search flag, no second API
# Search is enabled by default — always on, by design
response = client.chat.completions.create(
model="deepseek-v4-flash", # or qwen3.8-max, glm-53-flash, kimi-k3, minimax...
messages=[{"role": "user", "content": "What changed in EU AI regulation this week?"}]
)
print(response.choices[0].message.content)
That is the entire change. chat.completions.create works unchanged. The model receives current web context before it generates a reply.
What happens under the hood
| Step | What the gateway does |
|---|---|
| 1 | Detects the request needs current web context |
| 2 | Runs a web search over the open web (Tavily advanced) |
| 3 | Injects ranked web results into the model’s context window |
| 4 | Returns the model’s answer with sources embedded in the text |
The SDK, request shape, and response schema are identical to OpenAI. Your existing code does not change except api_key and base_url.
Why this is different from a Tavily wrapper
Most “web search with OpenAI” tutorials do this:
- Call Tavily to get results
- Manually build a prompt with the results
- Send the prompt to the model
- Hope the model cites the right sources
That is a RAG pipeline. It works, but it is 80–150 lines of code, a second API to provision, and a billing line for Tavily on top of the model.
Here the gateway does steps 1–3 for you. You send one message, one API call, one response.
Supported models
| Model | Notes |
|---|---|
| DeepSeek V4-Flash | Fast, low latency, strong reasoning |
| Qwen3.8 (Max / Plus / Flash) | Multilingual, long context |
| GLM-5.3-Flash | Chinese + English, cost efficient |
| Kimi K3 | Long document handling |
| MiniMax | Multimodal, code generation |
Switching models is changing one string. There is no separate gateway per model.
When to use this
- Your app already talks to OpenAI over the Chat Completions endpoint
- You need live web results without maintaining a search agent
- You want to compare DeepSeek, Qwen, GLM, and Kimi side by side
- You serve users outside China and need global web sources
- You want to pay with crypto, not a Tavily key + credit card
When not to use this
- You need structured, filtered results (price, citation URLs only) — use a direct Tavily call
- You need to display search results as cards (reviews, hotels, flights) — use a search UI API
- You need grounding citations as JSON arrays — the model returns them inline, not as structured fields
EasyRouterAI vs the alternatives
| EasyRouterAI | OpenAI native | OpenRouter | Tavily + any model | |
|---|---|---|---|---|
| SDK | OpenAI Chat Completions | Responses API | OpenAI-compatible | Custom |
| Tools required | None | tools=[{type:web_search}] | None | You build the loop |
| Search key | No (always on) | OpenAI plan or search scope | No | Tavily key |
| Search source | Global open web | OpenAI index | Provider dependent | Your Tavily index |
| Billing | Crypto or card | Card | Card | Tavily + model |
| Model choice | Multiple open weights | GPT-4o family | Many | Your weights |
FAQ
Do I need to pass tools=[{type:"web_search"}]?
No. Search is enabled by default and requires no tool declaration. The request is identical to a plain Chat Completions call.
Can I disable search?
No. Search is always on by design. If you need a non-search call, use a model that does not trigger web context, or contact support for a separate endpoint.
Is this really OpenAI-compatible?
Yes. The chat.completions.create endpoint, request body, and response schema are the same. Only base_url and api_key change.
Which search index is used?
The gateway runs Tavily advanced search over the global open web — English-language news, academic papers, official documentation, and primary sources. It is not limited to the model’s training region.
Which models are supported?
DeepSeek V4-Flash, Qwen3.8 (Max / Plus / Flash), GLM-5.3-Flash, Kimi K3, and MiniMax. More models are added as providers publish weights.
How do I pay?
In cryptocurrency (the supported option today is Bitcoin (BTC) on the network shown in your dashboard — no minimum top-up) or with a card on file. No separate search provider account is required.
Why not just use OpenAI’s own web search?
OpenAI’s web search needs the Responses API and an explicit tool on every request. If you already run a Chat Completions app, or you want DeepSeek/Qwen/GLM instead of GPT, this is the change you make.
Why model origin ≠ search source
Open-weight models from China are trained on large, multilingual corpora. That does not mean their search results have to be China-centric.
If you serve users in Europe, Southeast Asia, or the Americas, the model origin and the search source are two separate decisions:
- Model origin → reasoning quality, speed, language coverage, cost
- Search source → which web pages the model reads at request time
Why “Model From China” ≠ “Search Results From China” explains how decoupling them changes what your users see.
Related: Why “Model From China” ≠ “Search Results From China” Pay for LLM API with Bitcoin: A Developer Workflow Tavily Without a Second Bill