A capable general-purpose agent is not yet an engineering system. It still needs a place to work, useful instruments, and a trustworthy account of what happened.
For AI-assisted RFIC design, my preferred starting point is to reuse a general agent runtime and build the domain-specific laboratory around it. This is a statement about where to invest engineering effort, not a claim that all runtimes are interchangeable.
Separate the responsibilities
The word harness can become too broad to help with design. It may refer to the reasoning loop, tool definitions, an execution backend, a benchmark, or all of them together. I prefer to name the responsibilities explicitly.
This is a working decomposition, not a proposed universal software standard. The important property is that a change of model should not require a reinvention of the laboratory or a rewriting of the meaning of success.
Keep the inside expressive
An agent should be able to inspect raw diagnostics, write analysis code, explore authorized documentation, and compose ordinary engineering operations. A short list of rigid high-level commands can accidentally encode the developer’s preferred solution and exclude a better one.
That does not mean unrestricted access to everything. I want broad freedom inside a controlled workspace, with explicit authority at the boundary: which data may be read, which resources may be modified, what may leave the environment, and how much computation may be consumed.
The distinction matters. Restricting damage and expenditure is not the same as prescribing the next RF design decision. A runtime should enforce the former without quietly doing the latter.
Make feedback useful, not merely formal
A tool can return valid JSON and still be difficult to use. The result should identify what changed, what failed, which object was involved, and what the agent can inspect next. A failed geometric operation is a recoverable engineering event, not automatically a reason to terminate the task.
I want both structured results and access to the underlying evidence. The structured layer makes coordination reliable. Raw logs, native objects, and simulation files preserve the information needed to diagnose a surprising result. A summary should not be the only surviving record.
Verification is not self-assessment
The agent may propose that a design is ready. It should not have sole authority to define readiness, modify the acceptance checks, and certify its own result. Acceptance criteria, test conditions, budgets, and evidence provenance need an independent path.
For an RFIC task, that path can include design-rule checks, the appropriate connectivity or interface checks, electromagnetic analysis, and circuit-level evaluation of the resulting physical design. The precise checks depend on the task; a single generic “passed” flag is not enough.
Give the agent freedom to search. Give the environment responsibility for boundaries. Give the verifier authority over the claim.
Invest in what survives a model change
A reusable environment, an executable task definition, and an evidence-bearing artifact retain value when a stronger model arrives. A collection of special-case prompt patches may not. This is why I put the sandbox and the benchmark near the center of the work.
The benchmark should not become a second product larger than the design system. Begin with a few meaningful tasks, make their conditions explicit, preserve failures, and measure real costs. Expand only when the existing tasks stop answering the questions that matter.
My objective is a professional workbench that different agents can enter, use, and leave with a result someone else can inspect. A convincing transcript is not the deliverable. The engineering artifact is.
Ideas in progress. Corrections welcome.
Find me online