01 / ACTION

AI is being trained to work inside specialized business software, not just beside it.

OpenAI's October 6 research collaboration with Ironclad focuses on agents that can use contracting software, follow company rules, complete multi-step workflows, and check the result against the original request. Contracting is one example of a broader shift. As computer-use agents improve, more small-business tasks can move from generating a draft to entering, routing, updating, and completing work in the system where the official record lives.

02 / RULES

A request is not enough when the workflow contains business rules and exceptions.

The Ironclad work highlights a practical limitation: an agent can lose track of a rule during a long task. Most operating processes contain more rules than owners initially realize—approval limits, required fields, customer-specific terms, naming conventions, deadlines, and exceptions that experienced employees handle from memory. If those rules are not visible, the agent may produce a plausible result that is still wrong for the business.

03 / EVIDENCE

Reliable completion requires proof that each important condition was checked.

NIST's work on evaluation probes for agentic AI describes automated checks that compare outputs with trusted reference material and create audit trails linking decisions to supporting evidence. A small business does not need an advanced evaluation system to use the same principle. For consequential work, the final output should show which source was used, which rule was applied, what changed, and what still needs a person to approve.

THE CJC VIEW

Implementation beats experimentation.

The operating risk is not that an AI agent will always fail dramatically. It is that it will complete the wrong version of the task cleanly enough to look finished. A clear definition of done is therefore more valuable than a longer prompt.

Before an agent can safely act in a CRM, accounting system, contract platform, or project tool, the business should make four things explicit: the authorized scope, the required rules, the evidence of completion, and the conditions that stop the work for human review.

PRACTICAL NEXT STEP

Create an agent job card for one recurring task

  1. Choose one repeatable task that ends with information being saved, sent, approved, scheduled, or entered into a system of record.
  2. Write the trigger, required inputs, source of truth, and exact finished result in plain language on one page.
  3. List the five rules or checks an experienced employee applies, including the exceptions most likely to create a costly mistake.
  4. Specify the evidence the agent must return, such as source links, changed fields, calculations, timestamps, or a summary of actions taken.
  5. Name the conditions that require a stop and human approval, then test the card on five real examples before allowing any unattended action.

Agents will become more capable at using the software that runs a business. The companies that benefit will not be the ones with the most automation. They will be the ones that make good work visible enough to verify before they scale it.

Sources and further reading

  1. Advancing computer use with Ironclad — OpenAI
  2. Building Evaluation Probes into Agentic AI — National Institute of Standards and Technology
  3. Our framework for developing safe and trustworthy agents — Anthropic