This is what I wrote in my first essay about Grok Bot on Aug. 29, 2026, after my first 24 hours using it:
“As a working journalist, you might be wondering if I see any Grok Bot use cases for newsrooms? Easy answer: No. Maybe one day, but not right now.”
Twenty-four days into the experiment, my answer had changed:
“As a working journalist, you might be wondering if I see any Grok Bot use cases for newsrooms? Easy answer: Yes. Right now.”
So what changed?
For one, consumer AI is moving at light speed. Technologies such as mobile phones took decades to evolve from early cellular systems to modern smartphones. Consumer AI tools can change in meaningful ways in weeks — sometimes days.
But there is something more to it.
It wasn’t that Grok Bot suddenly became smarter. My impression of it changed because I stopped evaluating it abstractly as a journalism tool and gave it a real editorial task.
Specifically, I asked it to use judgment.
And handing over judgment to a machine is probably one of the scariest things a journalist can do. But in the age of AI, there is only one way to really find out what these tools can do: test them.
So I did.
Grok Bot performed so well in the experiment that I eventually decided to go public with the project.
Discovering that Grok Bot could search and sweep X was not terribly interesting in and of itself. What was interesting was discovering that, given a sufficiently clear set of editorial guidelines, it could make hundreds of small decisions about what mattered, what didn’t, what was credible, what was noise, what belonged on the beat and what should be left out.
And I also discovered where those rules broke down — and where the human had to intervene.
That was a key finding.
More on that in a minute.
I was intrigued by Grok Bot when it first came out in the second week of August and immediately dove in, wanting to be an early adopter. I created a couple of agents during my first 24 hours with it. That experience became my first essay about Grok Bot.
Initially, I wasn’t blown away. But I had an inclination that the first pass wasn’t the full picture, so I kept playing around with it.
At the same time, separate from Grok Bot, I was looking for a project to build in Codex. I had moved past vibe coding and wanted another project challenging enough to teach me something, but not so big that it would consume my life.
I do, after all, have a time-consuming full-time job as a journalist at the U.N., and this is an especially busy period as the race for the next U.N. Secretary-General heats up.
Every time I scrolled X, I was reminded of that race: candidates posting, diplomats posting, journalists posting, assorted commentary and plenty of noise.
So what to build?
Wait.
Grok Bot. Codex. X. U.N. Secretary-General race.
The light bulb went off.
I decided to combine all of the above into one project: a Grok Bot agent monitoring X for useful posts about the Secretary-General race, with the results eventually published on a site built with Codex.
I opened the Grok Bot app on my phone and created a new agent.
I named it The Stringer.
The Stringer would become part of a larger project I was building around the race: SG Candidate Watch, a broader site tracking the eight candidates, the selection process and where they stand on major global issues.
The Stringer became the project’s running X feed — the part designed to monitor campaign activity and relevant reporting in real time.
This essay isn’t about Codex versus Claude Code. My views on that I’ll save for a future essay. This essay is about discovering a credible use case for an autonomous agent and then building an editorial product around it.
Codex is how I packaged the experiment.
The experiment itself was designed to answer two questions:
Can an autonomous agent — in this case, The Stringer — turn the firehose of X activity around eight candidates running to lead the U.N. into a useful campaign diary and news feed for journalists?
And, more specifically:
Can I give an autonomous agent a reporter’s beat?
With AI, knowing what you don’t want to build is sometimes as important as knowing what you do want to build.
I did not want The Stringer simply reposting everything the candidates posted. I already followed all of them. That would get boring fast. I needed The Stringer to surface news around the world and show me X posts I would not otherwise have seen.
I also didn’t want it surfacing every random mention of the U.N. or Secretary-General race on X. That would quickly become overwhelmed by the very thing X produces in abundance: unverifiable noise and commentary meant for clicks.
And I didn’t want it sweeping the entire internet. I wanted to keep the experiment narrowly focused and lean into one of Grok Bot’s obvious strengths: X.
So I gave it editorial guidelines. I started with ten.
See the full standing rules used during the experiment.
Think of them as something close to an editorial or style guide in a newsroom. I told The Stringer what to include, but just as importantly, what to leave out.
If a candidate posts about what he or she ate for breakfast, skip it.
If a candidate meets with a foreign minister and discusses the Secretary-General race, that probably belongs.
I told it to look beyond the candidates themselves and sweep X for credible reporting, analysis and campaign activity, while excluding unverified accusations, inflammatory posts and other noise.
I also told it to treat every candidate the same.
Remain neutral.
Journalism 101.
Then I turned it loose.
I scheduled two sweeps — one in the morning and one at night — and closed my laptop.
The first results were impressive.
The Stringer showed me the posts it had surfaced and explained why they met the editorial guidelines.
But then it did something I hadn’t specifically asked it to do.
It showed me what it had rejected.
And it told me why.
That turned out to be one of the most useful parts of the entire experiment.
I wasn’t just looking at the finished output and thinking, this looks pretty good.
I could inspect its judgment.
I could see which posts it considered noise, which it believed were outside the scope of the beat, and how it was applying the editorial guidelines I had given it.
That is an important distinction.
I kept the experiment private for several days and let The Stringer run two sweeps a day. I reviewed what it surfaced and what it rejected. I looked for anything that crossed my editorial threshold, or anything useful it had consistently missed.
For the most part, I agreed with its decisions.
Dare I say, I agreed with its judgment.
That journalist’s fear of machine judgment started to wane.
I think we are on to something, I remember thinking to myself.
In fact, The Stringer was so cautious about following the editorial guidelines I gave it that I eventually loosened the collar a bit, allowing it to file posts like substantive analysis of the Secretary-General race from reputable NGOs and other credible sources.
The Stringer wasn’t just scraping posts.
It was making editorial selections inside the boundaries I had created.
And doing it quite well.
That is where my own 20-plus years of newsroom experience came into play. I wasn’t asking the agent to tell me what constituted good journalism.
I was giving it the standards, then checking whether it could apply them consistently.
So far, it could.
And this part of the experiment is not over.
The Stringer is still doing this today. By now it has run well over 100 sweeps, several times a day.
While I sleep, The Stringer sweeps.
When I wake up, I check the public feed and then open Grok Bot to see why it surfaced what it did — and why it rejected what it rejected.
Several of those sweeps have resulted in nothing being filed at all.
To me, that is significant.
An empty sweep is not a failure. It shows restraint. The Stringer scanned the beat, encountered plenty of material, and decided none of it cleared the editorial bar.
It wasn’t forcing itself to produce something simply because a sweep had run.
There were practical issues to consider.
At the U.N. there are six official languages — Arabic, Chinese, English, French, Russian and Spanish. And it’s an international campaign with candidates from various parts of the world who speak multiple languages. Limiting The Stringer to English-language posts made little sense.
I told it to include relevant posts in other languages and provide English summaries using the translation tools available through X.
It handled that without any problems.
Then there was the cost.
Each sweep required paid API usage. I had about $250 in xAI API credits available, but when I started the sweeps I had no idea how expensive this experiment would end up being.
Would all my credit get eaten up in a few days? Weeks? Months?
No idea.
There was only one way to find out.
Let it run and see what happens.
The first sweep cost about 28 cents.
How do I know?
The Stringer told me, as I had asked it to.
Fine.
I told The Stringer to keep future sweeps roughly in that range, and it generally did. Each sweep kept reporting back its cost, so I could keep track of what the experiment was actually costing me.
Useful.
But then came Grossi.
After a couple of days of testing, I noticed something strange.
The Stringer was surfacing almost nothing about Secretary-General candidate Rafael Grossi.
I quickly figured out why.
Grossi remains Director General of the International Atomic Energy Agency while also running for Secretary-General. Unlike the other candidates, a large portion of his public activity on X is therefore tied to his current job.
My original editorial guidelines told The Stringer to ignore candidates’ “day job” activity unless it was clearly connected to the Secretary-General race.
The Stringer was doing exactly what I had told it to do.
And that was the problem.
The Stringer wasn’t wrong. The rule was wrong for this particular case.
If Grossi traveled somewhere and met a foreign minister, was that simply IAEA business? Was it indirectly part of his campaign? What if neither the post nor the meeting made the distinction clear?
I was suddenly facing a real editorial judgment call.
So I reworked the editorial guidelines.
I told The Stringer it could include some of Grossi’s diplomatic meetings, even when they were not explicitly labeled as campaign activity, so long as they appeared potentially relevant.
But who determines what is potentially relevant?
The machine.
The agent.
The Stringer.
Judgment. There is that word again.
I let it run, closed my laptop and watched what happened.
The Stringer began surfacing some Grossi meetings at IAEA headquarters that were clearly part of his normal job.
That didn’t feel right either.
One of my core principles for the project was that the rules had to be applied as fairly as possible to all eight candidates. Giving Grossi extra exposure simply because he happened to remain in a high-profile international post would distort the feed.
So I adjusted the editorial guidelines again and told The Stringer:
Routine IAEA activity at headquarters should be left out.
Broader diplomatic activity away from obvious IAEA business?
Use your judgment.
And because this created a special case for one candidate, I added a Method Note to the site explaining the decision.
That episode ended up teaching me more about the experiment than almost anything else.
The machine had followed my rule correctly.
The real world produced an edge case.
The editor — me — changed the policy.
Then the editor watched what happened and changed it again.
That’s the human in the loop people talk about with AI, but in this case it wasn’t theoretical.
It was an actual editorial problem I had to solve.
Only after several days of private testing — and after I was satisfied that The Stringer was consistently following the editorial guidelines — did I move over to Codex and build the public site around it.
The Stringer didn’t eliminate editorial judgment.
It created a new place where editorial judgment had to happen.
Instead of personally deciding whether every individual post belonged on the feed, my role shifted upward: writing the guidelines, auditing the decisions, spotting anomalies, adjusting the policy and being transparent when an exception was necessary.
That, to me, is the most interesting part of the experiment.
And there was another lesson.
The clearest way I’ve found to see what works with AI is to give it a real task, with real standards, and see what happens.
About 24 days into this experiment, I had a very different answer from the one I had after the first 24 hours.
I had been trying to imagine a newsroom use case for a Grok Bot agent from the outside.
I couldn’t see one.
I had to put the thing on a beat to find it.
And the experiment is still going.
The Stringer remains an active feed, running several times a day, and I plan to keep it on the beat through the end of the Secretary-General race.
I still don’t know all the ways Grok Bot — or autonomous agents more broadly — will eventually be used in newsrooms. Those decisions will be made by journalists, editors and newsroom leaders as the technology develops.
But for the narrow task I gave it, The Stringer passed my editorial threshold.
A successful Grok Bot use case for newsrooms?
Easy answer: Yes.
It just took me 24 days, not 24 hours, to figure it out.