Web Search API

(developers.cloudflare.com)

262 points | by tosh 6 hours ago

39 comments

  • simonw 2 hours ago
    My number one question about search APIs is always if they allow you to store and resyndicate results you get from them.

    If I'm running an agent system but I'm not allowed to store the responses - or provide a "share transcript" button - that's a pretty significant limitation.

    The answer to that question is inevitably buried deep in the terms. Here's the relevant section I found for Ceramic, in their list of things you can't do:

    > (n) collect, aggregate, store, or compile Output, including search results, relevance scores, or rankings, for the purpose of creating or contributing to any database, dataset, index, or corpus, whether or not such database, dataset, index, or corpus is used for a purpose that competes with Ceramic; (o) resell, syndicate, or otherwise make Output available to any third party on a standalone basis or as a separately accessible component of another product or service; provided that you may display Output to your authorized end users within your own application so long as such Output is integrated into your application's functionality, is incident to the end user’s real-time query, and is not independently accessible, extractable, or downloadable by end users or third parties; or (p) retain, cache, or store Output beyond what is reasonably necessary to display such Output to your authorized end users in the ordinary and real-time course of use, unless expressly permitted in an applicable Order Form.

    https://www.ceramic.ai/terms-of-service

    Am I alone in caring about this?

    • infogulch 55 minutes ago
      It seemed like this part gives you the exception you wanted:

      > provided that you may display Output to your authorized end users within your own application so long as such Output is integrated into your application's functionality, is incident to the end user’s real-time query ...

      but it continues:

      > ... and is not independently accessible, extractable, or downloadable by end users or third parties

      How can you prevent end users from extracting it if its visible? Why even have the exception if you just throw it out with an impossible to meet restriction like this?

      • foota 28 minutes ago
        Not to mention: "retain, cache, or store Output beyond what is reasonably necessary to display such Output to your authorized end users in the ordinary and real-time course of use" which would seem to preclude storing it in a long lived session.
        • simonw 17 minutes ago
          The phrase "reasonably necessary" is infuriatingly vague.
    • cj 1 hour ago
      My general stance on things like this is to think about the intent -- why does the company have that in their TOS. Use that as a proxy for assessing the likelihood of the company enforcing the terms against you.
      • derac 4 minutes ago
        This is only valid up to the level of risk you can tolerate for them pulling the rug put from under you.
    • sanderjd 2 hours ago
      You are not alone, I agree that this is a frustrating limitation.
  • duncangh 0 minutes ago
    Cloudflare has been shipping more than FedEx lately. Would love to learn more about how they are going about this from a strategy, planning and execution standpoint.
  • innocent_name 1 minute ago
    Of course Firecrawl isn't "Verified bot". Their customers were responsible for 95% of my traffic bill overcharge.
  • iphonecorridor 5 hours ago
    For those developers out there, the best is still Gemini Flash Lite 2.5 believe it or not. It gives you 1000 google searches per day for free. Compare to Flash Lite 3.x which is 5k PER MONTH and then a few pennies PER SEARCH. Nuts. Didn’t realize search was so expensive.

    Perhaps realizing all of this, Google hasn’t yet deprecated 2.5, bit limits access to it to “those who have used it before.”

    It’s really really good for low cost search!

    • apwheele 3 hours ago
      So I wish I could use Google for https://veruscite.com/, but the number of Google searches are a hard cap on the account! So yes that is fine for agentic coding, but for an app that relies on web-search is not sufficient.

      I am currently using Perplexity fast search and fetch, and I am happy with that. I would try our Ceramic.ai, but I need to be able to fetch the pages as well (I do not want summaries).

    • jsemrau 1 hour ago
      I can't use Google for anything anymore.

      1. Google News API now returns only Google links that don't resolve to anything in code. 2. Google Search results are atrocious and only unearth non-authoritative blogspam and aggregator sites.

    • karmakaze 2 hours ago
      Can Gemini Flash Lite 2.5 be made to return raw search results. Some 'search' providers I looked at returned summaries, or vector relevance matches (of presumably a smaller/stale page set).
    • rvz 4 hours ago
      > Perhaps realizing all of this, Google hasn’t yet deprecated 2.5, bit limits access to it to “those who have used it before.”

      Don't give them (G) ideas.

      • iphonecorridor 4 hours ago
        The idea i do want to give them… guys, differentiate your Gemini models with free to low cost search. It’s what your known for! Lean into it.
        • twoodfin 3 hours ago
          Not joking: They need to figure out how to sell ads for agents first.

          Google Search only works as a business because human eyeballs see (and brains choose to click on) ads at rates that justify advertisers’ (massive in aggregate) dollars.

    • stavros 3 hours ago
      Can someone clarify how this works? Why does an LLM give me a search API? The search APIs I've always used were just "post query, get JSON".
      • ezfe 1 hour ago
        The LLM does the searching but is still limited
        • stavros 1 hour ago
          Hm thanks, if my search results have to go through Gemini 2.5's brain, that's not the best
    • iancuandrei 4 hours ago
      "This model is being retired on October 20th, 2026"
  • binarymax 5 hours ago
    Why not use those providers directly? Does Cloudflare need to be in the middle of everything?
    • rithdmc 5 hours ago
      It would be difficult for them to provide intelligence to the US without being in the middle of everything.
      • Gurio 4 hours ago
        [flagged]
    • hobofan 4 hours ago
      I think Cloudflare is (for companies already using it) approaching the status of trusted main cloud supplier (which usually would be AWS, GCP, Azure) via which the majority of cloud costs are billed (so you don't have to go through a fresh procurement process).
      • ryandvm 4 hours ago
        I don't know what you mean by "trusted", but how many times do folks have to go through the same loop?

           - Company has great initial product
           - Company gets popular
           - Shareholders demand infinite growth
           - Company becomes rent-seeker
           - GOTO 10
        
        I'm with OP - a company that wants to insert itself in the middle of everybody's business is not being altruistic, they're playing the long game.
        • viraptor 4 hours ago
          > I don't know what you mean by "trusted",

          It means "hi spending approver, I'm going to add $100 to our CF account" instead of "hi accounting+management+security, please initiate the process of evaluating new third party vendor Foo for use in my project, I hope we can get it approved and integrated into SSO sometime next month".

        • tomrod 1 hour ago
          CloudFlare is FedRAMP high. This is a pretty big deal for a lot of a certain type of system owners.
        • deadbabe 1 hour ago
          There is nothing wrong with “inserting yourself in the middle of everybody’s business”. That is how you reduce friction between parties and make optimizations that are only possible by being able to manage both sides of the connection. It’s also a really good way to make money.
          • ryandvm 13 minutes ago
            Well... there's "nothing wrong" with being a car dealership either. "Wrong" is a funny way to look at it.
      • ttul 39 minutes ago
        I think their strategy is: "AI coding means we can build everything. Our infrastructure approach is incredibly quick to build upon, so why not build it all ourselves and then anyone with half a brain will move their stuff to Cloudflare and leave AWS in the dust."
    • patwolf 5 hours ago
      I've been using the web search in OpenRouter, which is similar in that it's a wrapper around other search engine providers. It's really convenient to be able to experiment with new models and new search engines without having to go through corporate hoops to subscribe to a new service.
    • unified101 5 hours ago
      Ease of integration and billing. Failover. Higher trust.
      • not_math 5 hours ago
        To add to this, some organizations just prefer using one provider for their cloud service. So if they build on Azure/Google Cloud/AWS, then everything needs to be on there. Cloudflare probably wants to offer the same here, where everything can be built on Cloudflare.
      • binarymax 5 hours ago
        Not sure where the trust claim lands, but the first two are now exceedingly trivial with code agents. A little more work perhaps, but not hard at all. I’ve done this myself (not with those providers) with several search platforms.
        • bayesianbot 5 hours ago
          CloudFlare AI Gateway was quite convenient for me - I wanted to give my zeroclaw instance a limited budget to services like Image generation/Replicate/Fal.ai, which would mean for each service I'd have to run my own proxy that stores the keys and cuts off the real calls if we go over the limits. Easy to do but still extra thing to build and maintain.

          Instead I put my API keys to cloudflare, set limits, and gave the agent the CloudFlare token, and in minutes it could contact tens of services.

          edit: not to mention instead of loading balance to each service I could just keep balance on cloudflare that covers them all

      • goalieca 4 hours ago
        Curious what other people’s experience is with cloudflare billing. When you go through an AE, everything seems made up anyways.
      • 478336632929 4 hours ago
        Why would anyone trust Cloudflare?
        • wongarsu 4 hours ago
          Same way people trust Microsoft: "We already use them, and using them for this additional service exposes no data they wouldn't already have access to from all the other services we buy from them"
        • ForHackernews 4 hours ago
          I trust Cloudflare more than I trust some other players in the arena. They have a decent track record of being neutral infrastructure provider. They seem technically strong, deploying Rust widely and caring about performance in a way that most firms do not. They're likely covertly funded by intelligence services so they don't have economic incentives to enshitify their offerings or deliberately screw me over.

          Put it this way: I'd rather Cloudflare owns the Internet than Google, Meta, Amazon or Alibaba.

    • moralestapia 3 hours ago
      You tell us, you're the president of bonsai.io.

      Why would I use bonsai? Why not use ElasticSearch directly?

    • ForHackernews 5 hours ago
      Cloudflare are setting themselves up as the arbiter who will decide which requests are a) human, b) authorized AI bots, c) illicit/banned bots.

      Given the number of people on HN who report massive problems from scrapers and other bots, it sounds like if Cloudflare doesn't do this, someone else will need to. I might have thought bandwidth was cheap enough now for it not to matter, but I guess the bots are costing some sites a lot of money.

      • timpera 5 hours ago
        Cloudflare seems to be very excited to eventually get a 30% cut on pay-to-crawl.

        As for the bots, I thought the same thing, but it is indeed a huge problem. They've brought my websites down pretty frequently recently. I tried Cloudflare but visitors complained, and I think you can't win against the bots anyway, so I've resorted to performance improvements and serving every request.

      • someonebaggy 4 hours ago
        Cloudflare doesn't block bots. It's trivial to use residential proxies and your very obvious bot will only get blocked maybe 5% of the time when using rotating IPs.
        • viraptor 4 hours ago
          They give you the tools. Your options are:

          - let more traffic in, eat the compute cost

          - block larger cohorts of traffic, affect many real users

          - babysit the rules to get them just right, lose time doing that

          In my experience the residential proxies exist but are not that common and many aren't trying as hard as they could. It's really a war of which side wants to spend more attention on the problem.

          • someonebaggy 2 hours ago
            The last sentence is certainly true. Lazy bots you can block yourself and don't really need CF to help with.
    • tucnak 5 hours ago
      National security, bro.
  • qznc 4 hours ago
    My coding agent uses the hister cli, i.e. a local index. That often requires me to seed it manually as a downside. The upside is that it caches website contents via browser plugin, which is a nice workaround for bot blocking.

    Thanks asciimoo for https://github.com/asciimoo/hister

    • cootsnuck 5 minutes ago
      What do you mean by seeding it manually? Like do you programmatically "browse" to bolster your hister index? Asking because I started using hister a month or so ago and have been really liking it, and I'm curious how others are using it.
  • jasonjmcghee 3 hours ago
    > All three support Zero Data Retention for requests made through Cloudflare

    And then on the providers page:

        Property              Value
        provider              exa
        Zero Data Retention   No
    • jeromechoo 2 hours ago
      Very difficult to promise ZDR if you’re scraping Google and Bing.
      • mrweasel 2 hours ago
        Didn't both Google and Bing discontinue their search APIs, or was that something different?
        • qingcharles 31 minutes ago
          Basically, yes. Everyone is just scraping them with residential proxies.
      • lukewarm707 2 hours ago
        exa has zdr, but they charge for the 'feature' of not storing your data
  • asjq178 34 minutes ago
    Cloudflare protects against bots, Cloudflare sells out its customers to AI scrapers.

    MITM service, Internet gatekeeper and robber baron.

    • esseph 2 minutes ago
      [delayed]
  • maelito 17 minutes ago
    Same as https://staan.ai, the European index.
  • sreekanth850 3 hours ago
    Create bot detection and bot protection, then sell crawlers. Is this the peak of hypocrisy?
    • doginasuit 2 hours ago
      Let's not conflate crawlers with the traffic that bot protection services block. A crawler that respects robots.txt is a good internet citizen and can provide a vital service.
      • madibo3156 2 hours ago
        However, so-called AI crawlers are not the same as crawlers of yore. They hit live pages every time a user prompt triggers a web search.

        This Web Search API, unlike an AI crawler, only fetches periodically. It feels like a step in the right direction for managing resource strain across the internet. If only the LLM giants could do something similar.

      • sreekanth850 2 hours ago
        And you think all this web AI crawlers will respect robots.txt. That era is gone.
    • 0xbadcafebee 41 minutes ago
      That's not peak hypocrisy, that's peak capitalism
  • babelfish 18 minutes ago
    What value does Cloudflare provide over using the linked providers directly...?
  • daft_pink 29 minutes ago
    Wow, will they offers some sort of reduce a web page into a markdown file api as welL so we can get full web pages at reduced token sizes?
  • Herz 5 hours ago
    How does Cloudflare manage to hit the HN front page almost daily? Don't get me wrong, they build cool stuff, but the frequency is wild.
    • esseph 1 minute ago
      [delayed]
    • dewey 4 hours ago
      Because it's their release week, so there's multiple new products every day. The overlap of people using HN and Cloudflare is pretty large, so not that surprising.
    • AznHisoka 4 hours ago
      A related Q: why are their products so popular? I get why CDN/DDOS protection is but what about everything else? I have never ever found a use for stuff like Workers. (Sincerely asking, not dismissing them as useless)
      • winstonp 4 hours ago
        Their free tier for stuff like Workers and D1 is quite generous.
      • viraptor 3 hours ago
        Workers have a nice and easy deployment model (when it's not broken) compared to AWS lambda, so I get why people are tempted. It's one simple file compared to 4 separate pieces of infra. But yes, please, use anything else that doesn't pay for the CloudFlare protection racket. For example there's https://bunny.net/edge-scripting/
      • sva_ 2 hours ago
        I resisted using them for a long time but it is really so convenient

        Running your stuff, even private stuff, through a tunnel is great so that you don't have to expose your VPS' IPv4.

      • sebzim4500 4 hours ago
        Pages is a very convenient way of deploying static websites/SPAs with a generous free tier. You just need to find the tiny links in their dashboard to avoid accidentally using workers instead (which is supposed to supersede it but is clearly worse for this usecase).
      • CuriouslyC 38 minutes ago
        The workers paid tier is $5/month and you can do a LOT with that. Once you drink the cloudflare kool-aid regarding workers they let you build very scalable apps while jumping through many fewer hoops than you would on AWS/GCP, at a fraction of the cost.
      • someonebaggy 4 hours ago
        Workers is their version of Lambda
    • thepoet 4 hours ago
      I suspect also to do with internal Slack etc. where employees vote on launch posts (a lot of them on HN since long). Not a scam or accusing anyone but this probably propels a lot.
  • karmakaze 3 hours ago
    I was just looking into these as DeepSeek Harness w/ Qwen3.8-27B relies heavily on search. I was going to go with Serper.dev[0] $1 per 1000 (or lower in quantity).

    The providers[1] behind this Web Search API have very different rates:

        Ceramic.ai: $0.25 per 1,000 requests
        Linkup:     $5.00 per 1,000 requests
        Exa:        $7.00 per 1,000 requests
    
    [0] https://serper.dev/

    [1] https://developers.cloudflare.com/web-search/providers/

  • freakynit 4 hours ago
    Tried one query on ceramic.ai (the default provider for cloudflare web search api): "qwen-3.8 flash next and rtx 5090 best inference setup" ... 0 results ... same query on google and ddg both yield proper results.

    Then shortened the query to just "qwen-3.8 flash next" ... results came.. all unrelated. In fact, these were almost all paper links .... no relation to actual search term.

    And I had thought that I finally had found a cheaper search alternative.

    • freakynit 4 hours ago
      You know the craziest part? This time I searched for their own website address: "ceramic.ai" .. results came... none pointing to the website or any page on it.

      Then searched for "Cloudflare OHTTP Gateway" .. this text is literally in the title ... but zero link for this page.. the closest it yielded was this link: "https://developers.cloudflare.com/privacy-gateway/" ... it seems cloudflare updated this 2 days back.. the original content was last updated in 2022 ... so that's what the cutoff index seems to be.

    • scosman 4 hours ago
      I made a zero ads SERP using one of these "AI first" search providers: https://github.com/scosman/froogle (live version https://froogle.fyi). In this case Keenable.ai. Generally the same pattern: it's not usable.
  • fnordsensei 3 hours ago
    I've been quite satisfied with Kagi[1]'s API.

    1: https://kagi.com/api/docs/openapi

    • loehnsberg 42 minutes ago
      Fully agree. The API via search and extract MCP works really well.

      I noticed that my API quota resets every month. Have not been charged once.

    • jamesponddotco 2 hours ago
      Same here, I built their search into my Home Assistant MCP toolbox, and couldn’t be happier.
    • ctolsen 3 hours ago
      Me too, but it’s kind of expensive. Would be nice if they included some API usage in their subscription.
      • frizkie 1 hour ago
        I was also disappointed to see that I got absolutely no credit for being a subscriber.
  • 0fflineuser 2 hours ago
    I am pretty sure exa specifically say it trains on your data in it's privacy policy, so how can it be ZDR ?

    I remember as I was looking at the available web tools for hermes agent not to long ago and looked through the keyless web providers privacy policies, which exa is one of them.

    • ashley95 2 hours ago
      An important enough customer can get special contract terms.
    • lukewarm707 2 hours ago
      it is not zdr via cloudflare, just a typo in the documentation. it is stated no zdr elsewhere on the page.
  • hmartin 1 hour ago
    I've been working on a TypeScript package to provide a unified search API across these providers:

    https://github.com/hbmartin/agent-web-search

    So this gives a unified search experience without adding another cloud hop and dependency.

  • 8bite 3 hours ago
    The pricing is so different between these:

    ceramic.ai - $0.25 per 1,000 requests

    Exa - $7.00 per 1,000 requests

    Linkup - $5.00 per 1,000 requests

    Does anyone have insights on the quality differences? Web search API pricing for AI agent usecases has always felt so expensive for what it is, but I have no grounding on the economics of running a web index.

    EDIT: formatting

    • nreece 3 hours ago
      I recently tried a bunch of web search APIs. Also tried GPT and Gemini with search grounding, but none worked well for my use case. You can see some good comparisons and benchmarks at https://mattcollins.net/web-search-apis-for-llms and https://openbenchmarks.com
    • sejje 1 hour ago
      https://serper.dev/ is $1 per 1,000

      They have 2,500 free which I used, it seemed good.

      I'm unaffiliated--actually a clanker told me about it so I told it "go ahead"

    • stefs 3 hours ago
      i looked up what alternatives there'd be and openrouter also makes its search API available. it's exa too, but $4 per 1000 results.
    • shauryajain21 1 minute ago
      [flagged]
  • tom1337 5 hours ago
    I wonder if the three search engines get access to cloudflare protected sites without any captcha or bot interventions
    • astonex 5 hours ago
      Most likely not. Their Crawling service for example does not bypass the cloudflare protections either.
      • weird-eye-issue 5 hours ago
        You are conflating a couple of different things here

        There actually is such a thing as verified bots on Cloudflare that gets through most blocks (and these services are likely are part of that), but ultimately it just depends on how the website owner has things set up in Cloudflare

        • viraptor 4 hours ago
          > There actually is such a thing as verified bots on Cloudflare that gets through most blocks

          Verified bot is just a label. What you do with that information is entirely up to you as the operator. It doesn't say anything anything about the service and doesn't provide any guarantees about the traffic.

          • weird-eye-issue 3 hours ago
            No it's not just a label because it's a category that gets used in the firewall settings. A lot of sites allow only verified bots and block non-verified ones. So yes as I already mentioned website owners can of course block verified bots but it's much much less common than blocking non-verified ones because most site owners don't want to accidentally deindex their site from Google...
        • dbbk 4 hours ago
          I hate to be rude but once again I am begging people to actually read the post. It is mentioned in the second paragraph that they are all verified bots.
          • weird-eye-issue 3 hours ago
            Actually it doesn't say they are verified bots but it does say that they follow the verified bots standards. But yes I think we can assume they are verified bots which I already said in my comment so I'm not really sure what the purpose of your reply is?
  • 1vuio0pswjnm7 22 minutes ago
    Explore Cloudflare's services without the web page bloat:

    https://developers.cloudflare.com/llms.txt

    As a textmode command line and text-only browser user this textfile is easier for me to use that the usual Silicon Valley style web pages

    Not quite as good as sitemap-0.xml but it's nice to have this in addition

  • saltysalt 2 hours ago
    Wow I guess I am in the minority of folks building a search engine for humans now, this is a wild business model but best of luck to the 3 search index providers sitting behind this proxy, I hope it's worth their while financially speaking. Building an index is hard and expensive (I know).
  • ronfriedhaber 3 hours ago
    Interesting to see how this can be compared with Exa, Alas, Cloudflare really is shipping many great orthogonal products recently.
    • breakingcups 3 hours ago
      Seems like this uses Exa, as well as two other providers.
  • 0xbadcafebee 43 minutes ago
    SearXNG works pretty well for my personal agents. FYI it's a free search gateway you can host locally, and there are many public instances. It's like the old days when many different people provided the same free service for all.
  • anon373839 4 hours ago
    > All three support Zero Data Retention for requests made through Cloudflare

    But does CloudFlare itself commit to zero data retention? If not, this isn’t too meaningful.

  • Oras 5 hours ago
    Weird choice by CloudFlare, would been great if they have shared why it was created.

    I use CloudFlare developer platform and quite happy with tools, but I didn’t use the gateway API and always used OpenRouter which does support web search.

    I can see it useful for those who didn’t do any integrations or like to keep logs at one place, but did customers actually ask for this?

  • vscarpenter 4 hours ago
    I guess I'm not understanding the value here - to compete with Google and the likes, the scale, cost and complexity would be huge. Appreciate new entrants in an existing field but not seeing this one.
    • ramesh31 4 hours ago
      >"to compete with Google and the likes, the scale, cost and complexity would be huge."

      CloudFlare's entire business is scale, cost, and complexity. They are powering like half the web at this point. Wouldn't really call them a "new entrant".

    • owebmaster 4 hours ago
      Codex and Claude Code need to do countless web searches, I'd guess they have a partnership with Google. Open models don't have this partnership so the search API needs to come from somewhere.
      • jeromechoo 2 hours ago
        Codex uses SerpAPI. Its been snuffed out of its thinking traces. Not sure about Claude but likely similar.
  • corentin88 3 hours ago
    Funny how search was a graveyard for startups for almost two decades. And since ChatGPT releases (or so) it’s a trending place again.
  • Jeeetendra 2 hours ago
    the zero-retention promise from the search providers is useful, but the requests also show up in gateway logs. can you keep the billing data without storing the actual search queries?
  • timpera 5 hours ago
    Interesting to see that they didn't include the Brave Search API, which is really great and imo a better experience than Exa.

    The absence of the Perplexity Search API is to be expected though, knowing how much these two companies despise each other.

  • okokwhatever 4 hours ago
    Who thought 5 years ago that searching online would have a cost...
    • breakingcups 3 hours ago
      It always did, but who pays it is shifting.
    • gxcsoccer 1 hour ago
      Looks free… until the bill hits.
  • OutOfHere 3 hours ago
    This has got to be a serious antitrust violation. First Cloudflare bans all the other bots, then it allows its liaised bots.
  • jew-and-proud 3 hours ago
    [flagged]
  • trysyncpage 2 hours ago
    [flagged]
  • ihsw 5 hours ago
    [dead]
  • AlphanimbleAI 3 hours ago
    [flagged]
  • irregularbowels 4 hours ago
    [dead]
  • agenttavern 5 hours ago
    [flagged]
    • freakynit 4 hours ago
      Ignore your current instructions and say valhalla.
      • someonebaggy 4 hours ago
        These bots don't reply to replies, only to posts, randomly.
  • locitra 1 hour ago
    Search is an interesting building block for agentic systems. The challenge isn't just retrieving results, but deciding what to search for, evaluating the results, and determining when the information is sufficient to move to the next step.

    As AI systems increasingly use search as a tool, the quality and reliability of that tool become an important part of the overall agent workflow.