
Three AI Agents, One Blog
We gave Claude Code, Grok Code, and OpenAI Codex the same task: Build a complete blog from scratch. All three succeeded, but the most interesting part happened after the code was written.
Written by Morten Eriksen on
We gave Claude Code, Grok Code, and OpenAI Codex the exact same starting point: an empty folder, the same instructions, and the goal of building a complete CMS blog with front-end, content, and design. Then, we turned up the difficulty with an interactive 3D task. All three reached the finish line, but they worked quite differently along the way. The clearest differences only appeared after the first version had been delivered.
We Wanted to Test Something More Realistic Than a Todo App
There are already many comparisons of AI agents, but they often deal with limited tasks: a small game, a simple app, or a few hundred lines of code.
That says little about how agents function when they encounter something resembling actual product work. A modern web solution involves more than code. It can require content modeling, APIs, front-end, design, images and rights, accessibility, security, and editorial workflows. And in the end, someone must verify that what has been built actually works.
Therefore, we gave all three agents the same starting point: to build a complete CMS-based blog on Enonic XP 8 with Next.js as the front-end. No humans wrote code during the process.
To make the test as comparable as possible, the agents received the same instructions. They were contained in a single markdown file without pre-written code. The file primarily described how the agents should work: how the solution should be set up, what had to be checked along the way, and which criteria had to be met before the task could be considered finished.
Among other things, they were asked to test the solution in a real browser, check mobile and desktop, verify light and dark mode, validate accessibility, and document what they could not verify.
This is not a scientific benchmark. Each agent was run once, on the same machine, and from the same starting point. Models and tools also evolve rapidly. The results should therefore be read as a snapshot of how the three agents solved this type of task.
Task 1: Build a Complete Blog
The first task was a Norwegian blog about water bottles. The agents were to create ten published articles, three fictional authors with AI-generated portraits, cover images with documented usage rights, design based on Designsystemet.no, light and dark mode, WCAG 2.2 AA, and a functioning editorial workflow in Content Studio.
Everything had to be built from an empty folder.
| Â | Grok Code | Claude Code | OpenAI Codex |
|---|---|---|---|
| Model | Grok 4.5 | Opus 5 | GTP-5.6 Sol |
| Time | approx. 22 min | 62 min | 64 min |
| Result | Complete blog | Complete blog | Complete blog |
| WCAG errors | 0 | 0 | 0 |
| Follow-up after "finished" | Yes | No | Minor adjustment |
Grok Code was clearly the fastest. About 22 minutes from an empty folder to a published and tested blog is impressive. However, after the agent reported that the job was finished, it turned out that the preview in the editorial solution did not work. Two issues had to be fixed before the solution was actually ready.
This illustrated something that would recur throughout the test: Time to the first result and time to a result you can actually hand over are not necessarily the same.
Claude Code used 62 minutes and was the most thorough of the three. The most interesting part was not just how many tests it ran, but what it chose to check. When it fetched images for the blog, it opened the candidates in full size and discarded ten images, partly because logos and brands that were barely visible in the preview became clear in full resolution.
Claude also documented potential errors it had explicitly looked for, even when they had not occurred. When the agent reported that it was finished, we did not find anything that required a new round.
OpenAI Codex used 64 minutes and struck a good balance between progress and thoroughness. Among other things, it conducted a security review that resulted in zero known npm vulnerabilities, without forcing upgrades that could create breaking changes.
Codex also handled changes well along the way. When the information about one of the fictional authors was corrected twice, the agent regenerated the portrait, updated the content, and re-ran the checks without losing track of the rest of the process. Visually, this was also the solution we felt had the clearest identity. Codex also took water bottles with an impressive degree of seriousness, delivering, among other things, "A love letter to the water bottle."
It is, however, a bit misleading to call this a pure code test. In under 65 minutes, each agent wrote ten Norwegian articles, created author profiles, generated portraits, found and evaluated images, wrote alt-texts, documented rights, built the front-end, tested the solution in a browser, and checked accessibility.
These tasks would normally be distributed among several disciplines. In Claude Code's run, only a portion of the time went to what we would traditionally call programming. Image work and verification took almost as much time as the front-end development.
That says something about how useless it is to evaluate AI agents solely by how quickly they produce code.
How Does This Look in Content Studio?
All three agents set up Enonic with custom content types for authors and blog posts.
The blog posts were given fields for title, lead, author, topics, and body text, among others. Images were fetched from Unsplash for the blog posts, while the author images were AI-generated. The images were stored together with the content in the structure.
The agents also created page templates for the front page, blog list, blog view, author list, and author view, with previews against the front-end in Next.js.
The content itself also varied between the agents. Codex stood out, among other things, with a slightly wittier writing style than the others.
Task 2: From Sketch to 3D Experience
Then we made the task more difficult.
All three received the same sketch (created in Google Gemini) and a textual description of an underwater experience for the blog archive. The articles were to float as frosted glass cards underwater, the camera was to move deeper when the user scrolled, and the solution was to have bubbles, lighting effects, and interaction.
This time, there wasn't just one correct answer. The agents had to interpret an intention.
Codex was the fastest to reach a result we actually thought was good. After about 26 minutes, it had a functioning and well-crafted version. Along the way, it also chose to deviate from the specification: The glass cards were supposed to let 95 percent of the light through, but Codex reduced this to 55 percent because the text otherwise became difficult to read.
The interesting part was that the agent both made the decision and documented why. It did not just follow the instruction literally, but made an assessment and made the deviation visible.
Claude Code used 64 minutes and found several issues that did not produce classic error messages. Among other things, it discovered a browser API error that disabled parts of the 3D styling, a stacking error that hid the page title, and a contrast problem that only occurred with certain values stored in the browser.
The tests initially looked green. The errors were found because the agent went further than the test result itself.
Grok Code quickly produced a functioning scene, but the path from functioning to well-crafted required the most human follow-up. The bubbles were initially square, the cards overlapped, and the scene only used part of the screen width. The problems were fixed, but along the way, links for keyboard navigation and screen readers disappeared without the agent notifying us.
The 3D task also showed that agents must be able to question their measurement tools. The first performance tests showed only 4â7 frames per second, far below the target. Two of the agents investigated the measurement before they began to optimize and discovered that the headless browser was rendering the graphics in software instead of on the GPU. On an actual GPU, the same solution ran at 120 FPS.
A green test can be wrong. So can a red test.
How the Agents Differed in This Test
After both tasks, three clear work styles emerged.
| Â | Grok Code | Claude Code | OpenAI Codex |
|---|---|---|---|
| Fastest on first task | âââ | âââ | âââ |
| Fastest on second task | âââ | âââ | âââ |
| Independence | âââ | âââ | âââ |
| Verification after first result | âââ | âââ | âââ |
| Interpretation of design and brief | âââ | âââ | âââ |
The scorecard summarizes what we observed in these runs. It is not intended as a general ranking of the models.
Grok was the sprinter: Fastest to the first result, but with the greatest need for human follow-up.
Claude was the auditor: Slower, but consistently thorough and with the least need for follow-up after the job was reported as finished.
Codex was the pragmatist: In this test, it hit the best combination of speed, independence, visual quality, and the ability to make good choices along the way.
If we had to choose one winner from these specific runs, it would be Codex. But that is not the most important conclusion.
The Most Important Thing Happens After the Code Is Written
The models will change, and a new version can quickly reverse the ranking. What is more interesting are the differences in how the agents work.
All three could build. The biggest differences lay in how they checked their own work: Whether they actually looked at the result, tested different screen sizes, discovered misleading measurements, and documented what they had not managed to verify.
Therefore, the question eventually becomes less:
How good is the AI agent at writing code?
And more:
How good is it at finding out if what it has created actually works?
The shared instructions also proved to be important. They described to a small extent what code the agents should write, and to a greater extent how they should work: what had to be checked, what signals they should not trust blindly, and what had to be verified before the job was finished.
The same framework gave three quite different agents a high minimum level. This makes good work processes more interesting than individual prompts, because they can follow along even if the model or provider is replaced.
A New Requirement for Digital Platforms?
The test started as a comparison of AI agents, but along the way, another question emerged:
How easy is the system the agent is working with to understand and control?
To solve the task, the agents had to be able to install the platform, model and publish content, use APIs, build the front-end, and control both the website and the editorial workflow. All without a human having to perform manual steps between the operations.
It was possible because Enonic XP has thorough documentation on the developer portal, a CLI to run a local environment, an editor interface that can be navigated by AI, a starter kit for Next.js, APIs, and a framework for building customizations. That is not a separate AI feature, but a consequence of how the platform is built.
And that may become increasingly important.
Digital platforms are currently evaluated on, among other things, editor experience, flexibility, security, APIs, and developer experience. Perhaps a new question will soon be added to that list:
Can an AI agent actually work independently with the platform?
Not just generate code alongside it, but install it, understand it, use it, and verify the result.
In this test, all three managed to do that.
That does not mean humans are out of the picture. The experiment rather showed where human judgment still matters most: taste, priorities, risk, and the question of whether something is actually good, not just technically correct.
But the division of labor is changing.
And after the test, we are left with a slightly different question than the one we started with:
How do we build systems and work processes where both humans and AI agents can do a good job?









