If you already use the OpenAI Python SDK, you already know how to call a model:

from openai import OpenAI

client = OpenAI(
    api_key="sk-or-...",
    base_url="https://api.openai.com/v1"
)

response = client.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "What changed in EU AI regulation this week?"}]
)
print(response.choices[0].message.content)

Now add one thing: live web search. With most OpenAI-compatible gateways, you don’t need tools, a search key, or an agent loop. You only change the base_url.

OpenAI’s native web search requires the Responses API and an explicit tool declaration:

from openai import OpenAI

client = OpenAI(
    api_key="sk-...",
    base_url="https://api.openai.com/v1"
)

response = client.responses.create(
    model="gpt-4o",
    tools=[{"type": "web_search"}],
    input="What changed in EU AI regulation this week?"
)

print(response.output_text)

Two requirements follow:

  1. You must use the Responses API, not the classic Chat Completions endpoint.
  2. You must pass tools=[{"type": "web_search"}] on every request — there is no default.

What that costs and adds

  • A separate search:browse permission scope
  • A Tavily or Bing key if you are not on OpenAI’s plan
  • A max_search_results value you tune by hand
  • No upgrade path to other models without another SDK import

The one-line way

Point the same OpenAI SDK at a gateway that runs search automatically:

from openai import OpenAI

# Same OpenAI SDK you already use
client = OpenAI(
    api_key="your-easyrouterai-key",
    base_url="https://easyrouterai.com/v1"
)

# No tools parameter, no search flag, no second API
# Search is enabled by default — always on, by design
response = client.chat.completions.create(
    model="deepseek-v4-flash",   # or qwen3.8-max, glm-53-flash, kimi-k3, minimax...
    messages=[{"role": "user", "content": "What changed in EU AI regulation this week?"}]
)
print(response.choices[0].message.content)

That is the entire change. chat.completions.create works unchanged. The model receives current web context before it generates a reply.

What happens under the hood

StepWhat the gateway does
1Detects the request needs current web context
2Runs a web search over the open web (Tavily advanced)
3Injects ranked web results into the model’s context window
4Returns the model’s answer with sources embedded in the text

The SDK, request shape, and response schema are identical to OpenAI. Your existing code does not change except api_key and base_url.

Why this is different from a Tavily wrapper

Most “web search with OpenAI” tutorials do this:

  1. Call Tavily to get results
  2. Manually build a prompt with the results
  3. Send the prompt to the model
  4. Hope the model cites the right sources

That is a RAG pipeline. It works, but it is 80–150 lines of code, a second API to provision, and a billing line for Tavily on top of the model.

Here the gateway does steps 1–3 for you. You send one message, one API call, one response.

Supported models

ModelNotes
DeepSeek V4-FlashFast, low latency, strong reasoning
Qwen3.8 (Max / Plus / Flash)Multilingual, long context
GLM-5.3-FlashChinese + English, cost efficient
Kimi K3Long document handling
MiniMaxMultimodal, code generation

Switching models is changing one string. There is no separate gateway per model.

When to use this

  • Your app already talks to OpenAI over the Chat Completions endpoint
  • You need live web results without maintaining a search agent
  • You want to compare DeepSeek, Qwen, GLM, and Kimi side by side
  • You serve users outside China and need global web sources
  • You want to pay with crypto, not a Tavily key + credit card

When not to use this

  • You need structured, filtered results (price, citation URLs only) — use a direct Tavily call
  • You need to display search results as cards (reviews, hotels, flights) — use a search UI API
  • You need grounding citations as JSON arrays — the model returns them inline, not as structured fields

EasyRouterAI vs the alternatives

EasyRouterAIOpenAI nativeOpenRouterTavily + any model
SDKOpenAI Chat CompletionsResponses APIOpenAI-compatibleCustom
Tools requiredNonetools=[{type:web_search}]NoneYou build the loop
Search keyNo (always on)OpenAI plan or search scopeNoTavily key
Search sourceGlobal open webOpenAI indexProvider dependentYour Tavily index
BillingCrypto or cardCardCardTavily + model
Model choiceMultiple open weightsGPT-4o familyManyYour weights

FAQ

Do I need to pass tools=[{type:"web_search"}]?

No. Search is enabled by default and requires no tool declaration. The request is identical to a plain Chat Completions call.

Can I disable search?

No. Search is always on by design. If you need a non-search call, use a model that does not trigger web context, or contact support for a separate endpoint.

Is this really OpenAI-compatible?

Yes. The chat.completions.create endpoint, request body, and response schema are the same. Only base_url and api_key change.

Which search index is used?

The gateway runs Tavily advanced search over the global open web — English-language news, academic papers, official documentation, and primary sources. It is not limited to the model’s training region.

Which models are supported?

DeepSeek V4-Flash, Qwen3.8 (Max / Plus / Flash), GLM-5.3-Flash, Kimi K3, and MiniMax. More models are added as providers publish weights.

How do I pay?

In cryptocurrency (the supported option today is Bitcoin (BTC) on the network shown in your dashboard — no minimum top-up) or with a card on file. No separate search provider account is required.

Why not just use OpenAI’s own web search?

OpenAI’s web search needs the Responses API and an explicit tool on every request. If you already run a Chat Completions app, or you want DeepSeek/Qwen/GLM instead of GPT, this is the change you make.

Why model origin ≠ search source

Open-weight models from China are trained on large, multilingual corpora. That does not mean their search results have to be China-centric.

If you serve users in Europe, Southeast Asia, or the Americas, the model origin and the search source are two separate decisions:

  • Model origin → reasoning quality, speed, language coverage, cost
  • Search source → which web pages the model reads at request time

Why “Model From China” ≠ “Search Results From China” explains how decoupling them changes what your users see.


Related: Why “Model From China” ≠ “Search Results From China” Pay for LLM API with Bitcoin: A Developer Workflow Tavily Without a Second Bill