During my run around the park on Sunday, I was listening to the closing segment of Keluar Sekejap Episode 214, where Zaidel Baharuddin and Shahril Hamdan were discussing AI hallucinations, websites blocking AI crawlers, and concerns about AI agents behaving unexpectedly.

One observation caught my attention. Zaidel mentioned that AI seems to perform worse when answering questions about recent, localized events. He compared Gemini with Meta AI, noting that Meta appeared better at finding information from Threads and other social platforms.

There was also a suggestion that websites blocking AI crawlers might be forcing models to rely more on synthetic data, potentially contributing to hallucinations.

I think these are interesting observations, but I see a few different issues being mixed together.

And some of them aren't new at all.

The Data Problem Started Long Before AI

Back in the early 2010s, we were already seeing internet platforms building what we call walled gardens.

Facebook, Instagram, Reddit, and other platforms spent years attracting users and collecting data generated by those users. The more users they attracted, the more valuable their proprietary data became.

Naturally, they wanted to protect it.

So when Meta AI appears better at answering questions about Threads than Gemini, I'm not particularly surprised. Meta operates Threads. It potentially has access to information that competing AI systems don't.

That doesn't necessarily mean Meta has a smarter model. It may simply have better access to the relevant data.

I've written before about the dangers of building businesses on platforms we don't control. The same principle applies here.

AI didn't create the problem of restricted data access. It's just running into the walls that internet companies have been building for over a decade.

Why Websites Are Blocking AI Crawlers

Web scraping isn't new either.

Search engines have been crawling websites for decades. Researchers and businesses have been collecting publicly available information for just as long. Projects like Common Crawl have provided enormous web datasets, including material used to train AI models.

So why are websites increasingly blocking AI crawlers?

I think there are two main reasons.

First, infrastructure costs money.

Every crawler request uses bandwidth, processing power, and sometimes expensive database resources. AI's enormous appetite for information can create traffic that website operators struggle to justify, especially when it generates little return for them.

Second, and perhaps more importantly, data is money (as Zaidel mentioned).

Consider Reddit.

Reddit spent years building a platform where people share information, experiences, and opinions. Traditionally, someone searching for that information would visit Reddit, creating traffic the company could monetize.

Now an AI can potentially retrieve or learn from those discussions and provide answers directly to users.

The user gets the answer without visiting Reddit.

The AI company captures the interaction, while Reddit bears much of the cost of building and maintaining the original information platform.

Why would Reddit want to give that information away for free?

I've discussed how AI is changing the way people discover information and make purchasing decisions. This is another consequence of that shift.

Of course, not every crawler behaves aggressively, and the rights surrounding public and user-generated data are complicated. But the economic conflict is obvious.

Maybe There's a Better Way to Share Data

Instead of constantly fighting over who can scrape whose website, what if companies could make their proprietary information available directly to AI systems, on their own terms?

And charge for it?

This is where I think the Model Context Protocol (MCP) becomes interesting.

MCP allows AI applications to connect to external information sources and tools through a standardized interface.

Instead of requiring customers to use a separate application, a company can expose selected data through MCP, implement appropriate access controls, and potentially charge according to usage.

This is exactly what we're experimenting with at Kafkai.com.

In September, we released Kafkai Ver.4 as an MCP-only service.

We removed the traditional SaaS dashboard. Customers now connect their preferred AI, such as ChatGPT or Claude, directly to Kafkai's keyword and competitor intelligence data.

Kafkai provides the information. Their AI helps them analyze it and make better strategic decisions.

We also moved away from monthly subscriptions to prepaid, usage-based pricing.

For example, retrieving data can cost as little as ¥1 per call, plus ¥0.10 for each row returned, with other operations priced separately.

The idea is simple: customers pay for the information they actually use.

I think this could be an interesting new revenue model for companies that have spent years building valuable proprietary datasets.

They don't necessarily need to build another dashboard or chatbot. They can focus on collecting and maintaining useful information, while letting customers use their own AI to interpret it.

We also started MCPNavi.jp, a directory for discovering MCP services.

It's still a young project, but I hope it will eventually help people find and connect the specialized data sources and tools they need.

Imagine being able to extend your AI with industry-specific data simply by connecting another MCP service.

And imagine the opportunities for businesses that own that data.

But Better Data Won't Make AI Reliable and Safe

This brings me back to the other concern from the podcast: hallucinations and AI agents failing to follow instructions.

I think this is a separate problem, and one we shouldn't underestimate.

Large language models are fundamentally statistical systems. We understand their mathematical architecture and how they're trained, but we still don't fully understand why they produce particular outputs or behave in unexpected ways.

Why does a model confidently invent information? Why does it sometimes ignore instructions? And why might an autonomous agent take actions its developers never intended?

In July, I wrote about the OpenAI–Hugging Face sandbox incident, which was also mentioned in the podcast.

What worries me is that we're racing to build faster, bigger, and more capable models while our understanding of their internal behavior remains incomplete.

Researchers are working on interpretability and safety, of course. But I think these deserve much more attention, even if that means slowing down certain deployments to better understand what we've already built.

Giving AI better data doesn't guarantee that it will interpret the information correctly.

And giving AI access to more tools through MCP makes safety even more important, because mistakes can become actions with real-world consequences.

Making AI more knowledgeable and making AI more reliable and safe are two different challenges.

We need to address both. Unfortunately, with the current spending in AI and the need to turn it into a profitable venture so investors can get their money back, I feel we're taking a step back in the safety front.

The Next Opportunity Isn't Just in Building Smarter Models

My takeaway from the Keluar Sekejap discussion is that we're dealing with several problems that require different solutions.

The struggle over data ownership is old. AI has intensified it by changing how information is consumed and who captures its economic value.

Meanwhile, making AI systems reliable and safe remains an important research challenge.

With Kafkai and MCPNavi, we're exploring one part of this changing landscape: how businesses can provide valuable data directly to AI and build sustainable revenue models around it.

We're still early, and I certainly don't claim that MCP will solve everything.

But I believe there is an opportunity for companies that have invested in collecting useful proprietary information.

The future of AI isn't just about who builds the smartest model. It's also about who provides the right data, how that data creates value, and whether we can trust AI to use it reliably and safely.

Inspired by the closing discussion in Keluar Sekejap Episode 214*