Sooner or later — probably sooner — somebody will offer to make your newspaper archive “AI-ready.” The pitch deck will be full of glowing diagrams. The invoice will be painfully concrete.
Before signing it, ask one question: what will a machine be able to read afterward that it cannot read today?
For many publishers with open, crawlable article pages, the answer is: not much. Modern search and AI systems can already parse ordinary web pages. The real problem is not whether a model can enter the archive. It is whether the model — or a human reader — can identify what is in there, who reported it, when it happened and how one story connects to the next.
This is why the New York Post’s new AI product matters even to a local newspaper with no AI budget: it reveals that the quality of the answer begins with the quality of the archive.
What Hamilton actually does
On August 11, 2026, New York Post Media Group launched Hamilton inside the New York Post and California Post apps. Built with Google Cloud’s Gemini Enterprise Agent Platform, Hamilton combines conversational search, personalized news briefings, recommendations and notifications.
The conversational search is the part local publishers should study. According to A Media Operator’s reporting, it searches five years of the publisher’s archive at launch, answers questions in plain language and cites the Post articles behind each answer. Newly published stories can be ingested and made searchable within seconds, Axios reported.
The boundary is deliberately narrow. New York Post CTO Ariscielle Novicio described the rule given to the system this way: “If we don’t write about it, do not even try to answer.”
Hamilton does generate conversational answers and personalized digests, so it would be inaccurate to say the software writes nothing. What it does not do is original reporting or independent editorial judgment. The Post controls the source material and editorial boundaries; the model retrieves, organizes and presents journalism the newsroom already produced.
Even one of Hamilton’s least glamorous components makes the same point. The Post built a caching layer so thousands of similar reader questions do not create thousands of fresh, paid model calls. The visible product is AI. The work underneath it is information architecture, cost control and editorial discipline.
Readable is not the same as useful
There is an important qualifier to the claim that an archive is already AI-readable. Its pages must actually be accessible under the publisher’s chosen terms. A broken URL, an accidental noindex, a login wall or content rendered in a way a crawler cannot retrieve is still a locked door.
“Readable” also does not mean “free to train on,” “licensed for reuse” or “available to every model.” Access, licensing and technical readability are different questions, and publishers should decide each of them deliberately.
Once a system has legitimate access, however, another problem appears. What happens when almost everything has been filed under one category called “News”? What can it establish about provenance when the byline on thousands of stories says only “Staff”? What does it learn from a photograph with no caption, no alt text and a filename such as IMG_4827.jpg?
At 4media, this is a recurring pattern during newspaper migrations. Inside the disorder may sit 20 or 30 years of the only continuous record of a town: elections, school boards, floods, business openings, obituaries, high-school championships and the people who shaped the community.
Usually, nobody made one disastrous decision. The disorder accumulated one deadline at a time. A broad category was faster than choosing the right one. “Staff” was easier than maintaining author accounts. Captions and tags felt optional because the article still looked fine when it went live. Two decades later, those shortcuts have become the architecture of the archive.
An AI interface does not repair that history by appearing above it. It exposes the condition the archive was already in.
Google is not asking for a second, AI-only website
Google’s current guidance for generative AI features in Search is unusually direct: the foundations of visibility in AI search are still the foundations of SEO. Google recommends valuable, non-commodity content, a clear technical structure and pages its systems can crawl and understand. It says there is no special schema.org markup required for generative AI search and advises publishers to evaluate supposed AEO or GEO services against the same established SEO principles.
That matters because “AI readiness” is already becoming a label that can be attached to almost any service. A new text file, a block of machine-generated FAQs or a promise to “chunk” every article for an LLM may sound advanced without improving the underlying journalism at all. Our analysis of whether AI systems actually use llms.txt found the same gap between adoption and evidence.
Seven fields that decide what the AI finds
The durable work is less fashionable:
- Specific categories and consistent tags. “News” is a container, not a useful description. Geography, institution, beat and recurring topic labels give retrieval systems meaningful ways to narrow a search.
- Named authors with real profile pages. Google’s people-first guidance encourages clear bylines and information about who created the content. E-E-A-T is not a single ranking factor or score. It is Google’s framework for thinking about experience, expertise, authoritativeness and — most importantly — trust. CMS4media’s Authors module connects a byline to a profile, biography and the writer’s other work.
- Reliable publication and modification dates. An AI assistant trying to reconstruct a mayoral race or storm response needs to understand the sequence of events. Silent date changes and missing update notes turn a timeline into guesswork.
- Captions and descriptive alt text. A photograph is part of the reporting. A vision model may be able to guess what an old image shows, but it should not have to guess who is pictured, where the photo was taken or why it matters. Captions preserve editorial context; alt text makes the page more accessible.
- Stable URLs, canonical signals and honest redirects. A citation is only valuable if it still resolves to the story it names. Our guide to redirects and HTTP status codes for news archives explains why sending every retired article to the homepage creates confusion instead of preserving value.
- Article-level structured data. Google says Article and NewsArticle markup can clarify a story’s headline, author, images and publication dates. It is not a magic ranking switch and not special AI markup. It is a consistent description of what is already on the page.
- Source relationships and internal links. A developing story is rarely one URL. Links between the first report, public records, corrections, follow-ups and analysis give both readers and retrieval systems the surrounding record.
None of these fields can turn thin reporting into authoritative journalism. They do something more practical: they stop authoritative journalism from disappearing inside its own database.

The archive should work before the chatbot arrives
Hamilton is an audience product, not merely an archive demo. The New York Post team says it is moving beyond raw pageviews to measures such as visit frequency, session depth, retention, conversion and reader lifetime value. That is the right test: not whether an AI answer looks futuristic, but whether it helps readers build a deeper relationship with the publication.
A local publisher can start extracting the same kind of value without building a conversational assistant. Old reporting should appear beside current reporting. A reader following a new zoning dispute should be able to reach the original development proposal, earlier council votes and the reporting on the officials involved.
The Related Articles module in CMS4media, for example, can resurface stories by category or tag. That helps a human reader today and rewards the same classification work a future retrieval system will depend on tomorrow.
The migration is another decisive moment. In a recent CMS4media project, Hartmann Media Consulting moved six local news portals. On its larger sites, the work included years of articles, photographs and metadata, followed by redirect cleanup, indexing checks and Search Console monitoring. An independent audit cited in our migration case study found that organic search traffic held steady while Google Discover grew by close to 41 percent and Google News by 14.9 percent.
This is what responsible archive work looks like: preserve what has value, map the old URLs, expose technical problems and measure what happens after launch.
No responsible CMS can reconstruct a missing byline or invent a reliable caption for a photograph whose context has been lost. What it can do is preserve the information that still exists and make better habits easier from this point forward.
CMS4media’s AI Assistant, for example, can help editors fill repetitive fields such as titles, introductions and SEO descriptions from the reporting they have already entered. It can reduce the daily friction that caused metadata to be skipped in the first place. The newsroom still reviews the result. AI should help maintain editorial structure, not substitute for it.
Ask vendors three questions, not one
Soon every CMS vendor will say it has AI. Most will be telling the truth. The label will therefore tell publishers almost nothing.
Ask instead:
- What have you already built and migrated? Look for live publishers, intact archives, working author records, stable URLs and measurable post-migration results.
- What are you building now? A roadmap should solve problems editors and publishers can recognize, not simply follow the vocabulary of the latest keynote.
- How is the next roadmap decided? At 4media, regular conversations with publishers and patterns in real support requests shape what we build next. That process matters more than the number of AI logos on a sales slide.
Your archive may need cleanup. It may need a careful migration. It may need better taxonomy, author records, captions, redirects and structured data. Those are real projects, and some require serious work.
What it probably does not need is an expensive costume called “AI readiness.”
The New York Post did not make five years of journalism valuable by placing Hamilton in front of it. Hamilton made the existing value easier for readers to reach. For a local publisher, the order of operations should be the same: protect the reporting, organize the record and make every source easy to retrieve. Then decide which new door you want to build into it.
Metadata will never be the most exciting line in a press release. It still runs the show.
