Skip to content

Shipping a Feature in Python and Prose

Some agent capabilities are maintained across executable implementation and model-facing instructions, making cross-representation coherence part of correctness even though current evidence does not establish the defect rate or the best way to enforce that coherence.

On March 19, a pull request landed in Weaviate’s open-source skills for coding agents. Its headline feature was straightforward: PDF import.

The implementation lived where a programmer would expect. A Python script learned to turn PDF pages into images, create a multimodal collection, handle other input formats differently, cast values against a schema, and accept new command-line options.

But the feature was not finished there.

The same pull request changed SKILL.md, the Markdown file that tells an agent what the skill can do. It added PDF to the supported import formats. A longer Markdown reference acquired the new invocation forms, flags, PDF collection rules, and revised type-conversion guidance. The merged change put one feature into two kinds of artifact at once.

A file tree encourages a familiar distinction: Python is behavior; Markdown is documentation. Here, that distinction is incomplete.

The Python determines what the tool can do. The Markdown helps determine what the model will try to make it do.

One feature, two interpreters

Claude Code’s plugin documentation describes skills as capabilities the model can select for a task. A skill’s SKILL.md supplies instructions; scripts and other resources can sit beside it. The two files are not executable in the same sense. Python runs under a conventional interpreter. The instructions are read by a probabilistic model.

But both can participate in the same operation.

Imagine the Weaviate script gaining PDF support while its instructions still advertise only CSV, JSON, and JSONL. The capability exists, but the model-facing description of it is stale. Reverse the mismatch: instructions describe a command option the implementation no longer provides. The prose can remain valid Markdown while becoming wrong operating guidance.

Neither hypothetical failure was demonstrated by the March pull request. The maintainers changed the files together; they did not deliberately leave one stale and measure the result. What the episode establishes is narrower: a real feature was represented simultaneously in executable implementation and instructions consumed during agent operation.

An August 28 study asks how often that kind of pairing appears.

Ahmed Hereiz and four colleagues examined 1,926 repositories, 8,351 Claude Code plugins, and 77,773 plugin-touching commits. The paper is an arXiv v1 manuscript submitted to ACM, not yet a completed peer-reviewed publication. Most plugin components did not show above-chance coupling with one another. Inside skills, scripts and supporting Markdown did.

The researchers began with 13,587 pull requests touching skills. They found 1,908 in which a script and Markdown file changed inside the same skill, then narrowed that set to 323 in which both were modifications rather than newly introduced files. From those 323 they drew a stratified sample of 64 pull requests for human coding.

Fifty of the 64 were judged functionally coupled. Fourteen were not. Two raters reached Cohen’s kappa of 0.74 before resolving 13 disagreements.

That is the source of the paper’s 78 percent figure. Its denominator changes its meaning completely.

It does not say that 78 percent of plugin changes require a code-and-prose update. It says that among a sample already conditioned on scripts and Markdown having changed together, the researchers usually found a shared interface, behavior, value, or repository change connecting them.

The finding is evidence of recurring co-maintenance. It is not a production defect rate.

The old problem and the new consumer

Software engineering has had a version of this problem for decades.

Research on code-comment co-evolution long predates coding agents. Source changes can leave comments describing an interface or behavior that no longer exists. The program may still compile; the developer reading it receives the stale part.

LLM applications moved prose closer to the behavioral path before agent plugins arrived. A 2024 study of 1,262 prompt changes across 243 repositories found prompts evolving during feature work and sometimes developing logical inconsistencies or prompt-output misalignment. PromptPex later treated prompts as code-like artifacts whose modifications can introduce regressions, generating tests intended to expose them.

So “prose is software now” is not the discovery. It is too broad, and it is not new.

What changes in an agent skill is the consumer.

A stale comment usually misinforms a developer trying to understand a program. A stale skill instruction can misinform the system choosing an operation. The old maintenance failure sits between software and its human explanation. The newer one can sit inside the path from a task to an action.

That difference is enough to move the practical boundary without pretending the problem has no ancestor.

A file extension is not a risk classification

The Hereiz paper also reports that many commits labeled as documentation were doing work its diff-level classifier treated as feature or bug maintenance.

Those exact percentages deserve caution. The procedure combines human judgment and LLM-assisted classification, and the authors list ambiguous cases, model sensitivity, and subjective coding decisions among their validity threats. The reclassification is a study result, not an audited definition of what docs means everywhere.

The simpler point survives without the headline percentages: a .md extension tells you syntax, not runtime role.

Markdown can be a README for a human, instructions selected by an agent, a specification consumed by a generator, examples copied into a prompt, or several of those at once. The extension alone cannot tell a reviewer whether changing the file can alter behavior.

The same problem runs in the other direction. A Python change can be locally correct while leaving the model-facing description of its interface stale. Static tooling has no universal edge between a sentence describing an option and the parser that implements it.

The useful object is therefore not “documentation” or “code.” It is the shared claim represented in both.

The Weaviate change contains several such claims: PDF is accepted; particular command forms invoke it; certain options exist; PDF import follows particular collection rules. Each fact appears in implementation and in agent-facing instructions. The maintenance question is whether those representations still agree after either one changes.

Behavior-bearing does not mean beneficial

Calling instructions part of the behavioral surface creates another temptation: assuming more instructions must be better.

The evidence does not support that either.

A February study of repository context files such as AGENTS.md found that coding agents generally respected the added instructions and explored repositories more broadly, yet task success tended to fall while inference cost rose by more than 20 percent. A July ablation used two agents on 17 real tasks across three repositories, totaling 288 evaluated runs, and found no measurable correctness effect from context-file strategy within the range it tested.

Those studies are not experiments on Claude Code skills. They cannot tell us whether a stale SKILL.md breaks a paired script. They establish a different boundary: instructions can affect an agent without helping it.

So the maintenance objective cannot be “preserve every instruction.” An instruction may be obsolete, redundant, overly broad, or unnecessary. Sometimes the right repair is deletion. Sometimes a schema or generator can remove duplicated facts. Sometimes a behavioral test can detect the mismatch more directly than an explicit relationship between files.

The Hereiz study does not compare those interventions. It also does not establish how often one-sided changes cause user-visible failures. Co-change history can reveal a relationship without proving that every paired edit was necessary.

What actually moved

The study is also a snapshot of a young, narrow population. Its data were collected in April 2026. The median repository age was 80 days. The baseline dataset excluded repositories with fewer than 10 GitHub stars. The authors explicitly warn that the results have not been shown to generalize to Cursor, Copilot, Gemini, or AI-native software as a whole.

Within that boundary, the strongest conclusion is smaller than “documentation is code” but more useful.

Some software features are now written twice for two different interpreters.

One representation is consumed by an ordinary runtime. Another is consumed by a model deciding what operation to attempt. The representations can use different languages, live in different files, pass different validation checks, and still describe the same capability.

When that happens, correctness can become a property of the agreement between them.

The March Weaviate change makes the shift concrete. PDF support was not just the code that made import possible. It also included the instructions that told an agent PDF import existed and how to use it.

The important boundary did not move from code to prose.

It moved from the individual file to the fact both files were trying to represent.

Sources