Catalyst + 3rd Party Modules

On release, what options will module authors have for integrating with Catalyst’s features?

(IDK how to talk about this yet, so I’m going to refer to Catalyst as an “actor” lol)

Specific questions:

  1. Will Catalyst know about 3rd party Perspective components?
  2. Will Catalyst know about my expression functions, scripting functions, and custom binding types?
  3. Can Catalyst be integrated with custom Designer workspaces? (I.e. can I add a Catalyst pop-up button to my Client Resources editor?)
  4. WRT all of these, how can additional context be included? Like if I wanted to give the model usage examples?

I just did an internal POC for adding Catalyst's features to SQL Bridge, which is basically what you'd do as a third party module author.

You'd declare a soft dependency on Catalyst (so that your module still works if users don't install Catalyst), and if Catalyst is available provide information to the ai context it provides.

@paul-griffith do you know if your team is planning on having an SDK example up in the beta time frame? Or did I just volunteer myself for that? :winking_face_with_tongue:

Disclaimer: Exact method names, class names, etc, subject to change.

Yes, directly; the component assistant walks the schema that's already provided for all components today.

Yes, expression and scripting functions are looked up directly via tool calls by the running agent; if it's something you can enter as a user it's provided to the agent.
Custom binding types - out of the box not exactly. The current plan is to provide an optional extension hook on the binding registry. You can either provide your own prose framing (and hope for the best) or you can fully opt in to a schema-based binding (and action, and so on) approach. You declare a json schema ahead of time, it's enforced by the backend, and it's used to validate whatever the model gives us deterministically. This is opt-in so if you don't care about Catalyst you don't have to think about it, but all of our first party stuff will obviously participate.

Yes. In both the designer and the gateway, the platform context has a dangling "AiContext" attached to it that's basically "supplied" by Catalyst. In the designer, you have a few paths to use this context:

A commented, cut down snippet from the script console assistant:


private @Nullable JComponent buildAiPanel() {
aiSession = context.getAiContext()
  .map(ai -> ai.assistant()
    // "Creating an agent" is just passing in markdown and a set of tools
    .agent(() -> createAgent(ai))
    // EmbedHost is the wrapper around the actual Swing surface area you're interacting with
    .host(EmbedHost.forTextEditor(TextlikeComponent.of(buffer)))
    .promptAssists(
      // use any of our "standard" prompt assists, or plug in your own
      ai.promptAssists().tags(),
      ai.promptAssists().systemLibrary(),
      ai.promptAssists().userLibrary())
  )
  // Either mount it as a popup or "inline" if you've got a place for it in your UI
  .flatMap(EmbeddedAssistantBuilder::inline)
  // If the user doesn't have the module installed, *or* hasn't configured a model provider yet in the designer,
  // this will be empty
  .orElse(null);

return aiSession == null ? null : aiSession.component();
}

private @NonNull Agent createAgent(DesignerAiContext ai) {
  return ai.agents()
      .scriptingAgent() // Use our base scripting framing
      .framing(
          """
          ... prose ... 
          """)
      .tools(() -> Tools.fromAnnotated(new ConsoleInspectionTool(console)))
      .build();
}

Specifically for e.g. components, bindings, actions in Perspective - either put them into your schema description or in the additional AI specific "framing" you get to add.

On custom workspaces (really any embed you add to the designer) you can create your "agent" with whatever arbitrary context you want.

I guess I'll also add for posterity, DesignerAiContext extends the base AiContext which provides the following, useful if you want to do your own inference of any kind outside of our assistant framings:

Awesome, this gives me a bunch to start thinking about. :grin:

In this example, is this the section that defines the target for AI interactions?

EmbedHost.forTextEditor(TextlikeComponent.of(buffer))

I guess I'm most interested ATM about how to hook this whole thing into existing UI's.

So, yes. There's a slightly convoluted stack of things here, because of threading/presentation/separation of concerns reasons. We've also debated this architecture a fair bit and may still change some parts of it.

EmbedHost is just a carrier for a PromptComposer and a ReplyRendererFactory.
PromptComposer is given the user's actual prompt input and any @ references, but doesn't return those - it uses those to generate its own text output format. So for the basic TextEditorPromptComposer, you return a string that looks like:

<editor>
...full text...
</editor>
<selection>
...user selection...
</selection>

If you have a more complex UI, you would provide your own prompt composer that passes the state of those extra fields in.

Then the ReplyRendererFactory is two callbacks, one for the "statement" branch and one for the "suggestion" branch.
The format is prose only, i.e. the end user asked "what is this" and got a response from the model. The easiest thing to do is display this in a com.inductiveautomation.ignition.designer.gui.MarkdownPanel.
The suggestion branch is more interesting - you get a prose description from the model and (usually) a JSON string in value. You can deserialize that into a real object, run validation on it, and turn that into a series of actual actions to take in the designer. Then you return a ReplyView, which is just a rendering JComponent and a list of buttons for us to attach to it. Most likely, this will be just 'Apply', but could be something like "Insert" vs "Replace", or whatever else you want it to be.

record EmbedHost(PromptComposer composer, ReplyRendererFactory rendererFactory)

@FunctionalInterface
public interface PromptComposer {

  /**
   * Builds the host-context blocks for one user submission. The framework appends the canonical
   * references block and the user's prompt after whatever this returns.
   */
  CompletableFuture<String> compose(String userPrompt, List<ContextReference> references);
}

@FunctionalInterface
public interface ReplyRendererFactory {

  /**
   * @param applied runnable an apply action invokes <em>after</em> its host write-back returns.
   *     Mounters (e.g. the popup-anchored trigger) listen for this signal to dismiss themselves;
   *     inline mounters are free to ignore it.
   * @return the renderer that handles every reply in this session.
   */
  ReplyRenderer create(Runnable applied);
}


/**
 * Renders one model reply into a Swing view with optional apply actions. One renderer is held per
 * assist session and called on each incoming reply.
 *
 * <p>Both methods run on the EDT and may construct fresh Swing components on every call; the panel
 * installs the returned {@link ReplyView} into the conversation column. The host picks which method
 * to call based on the discriminator in the structured reply the harness produced: informational
 * replies route to {@link #renderStatement}, apply-able replies route to {@link #renderSuggestion}.
 *
 */
public interface ReplyRenderer {

  /**
   * Render an informational reply with no apply-able actions.
   */
  ReplyView renderStatement(String description);

  /**
   * Render an apply-able suggestion. The {@code value} is the literal content the user can write
   * back to the host editor; {@code description} is the prose framing. The returned view typically
   * carries one or more apply buttons (Insert, Replace, …).
   */
  ReplyView renderSuggestion(String description, String value);

}

record ReplyView(JComponent component, List<JButton> actions)

That’s pretty cool, I get the vision now. It sounds quite flexible.