Skip to main content

Embedding an agent in a web page

Put an agent on your own site — a marketing page, a docs page, a help centre — where visitors can talk to it without signing in to anything. Your site stays static: there is no proxy to deploy, no server-side secret, and no CORS grant to arrange, because the chat itself runs on Delegate's origin inside an iframe.

Two things make that safe to do with a token pasted into public HTML:

  • A share link is scoped to exactly one agent, is chat-only, and can be revoked instantly. It cannot list agents, provision anything, read tool credentials, or change the agent's configuration.
  • An embed allow-list on the link names the sites permitted to frame it.

This is a different credential from the service-user key described in the security model. That key is a master key for your tenant and still belongs only on your backend. A share-link token is not, and is meant to be published.

In the Delegate UI: Agent → Share Links → Create link. Set Embed on to the sites that may embed it.

Or over the API:

POST /v1/agent/{agent_id}/share-links
X-Api-Key: <a user API key belonging to an admin of the tenant>
{
"label": "verinfast.com marketing site",
"expires_in_days": 365,
"allowed_origins": ["https://verinfast.com", "https://*.verinfast.com"]
}
{
"link": { "id": "…", "token_prefix": "dsl_a1b2c3d4", "allowed_origins": ["…"] },
"share_token": "dsl_…",
"share_path": "/share/dsl_…",
"embed_path": "/v1/share/dsl_…/embed",
"embed_script_path": "/v1/share/embed.js"
}

The token is returned once. Only its hash is stored, so a lost token means minting a new one.

Use a user API key, not a tenant service key. Managing share links requires an identified tenant admin; a tenant-level key authenticates as no particular user and is rejected with a 403.

2. Drop in the widget​

One script tag, no build step, nothing to import:

<script src="https://agent-api.verinfast.com/v1/share/embed.js"
data-token="dsl_…"></script>

That gives you a launcher button in the bottom-right corner that opens a chat panel. The iframe is created on first open, so a page nobody chats on costs one small script and no agent request.

To place the chat inline instead — a dedicated page where the chat is the content — name a container:

<div id="agent" style="height: 560px"></div>
<script src="https://agent-api.verinfast.com/v1/share/embed.js"
data-token="dsl_…"
data-mount="#agent"></script>

data-mount present means inline; absent means floating launcher.

Layout attributes​

AttributeDefaultApplies to
data-tokenrequiredboth
data-mount—inline mode, a CSS selector
data-positionrightlauncher (left or right)
data-offset20pxlauncher
data-width380pxpanel
data-height560pxpanel, and inline if the container has no height
data-z-index2147483000launcher and panel
data-openfalsetrue opens the panel on load
data-launcher-labelChat with uslauncher's accessible label
data-launcher-icon💬launcher's glyph

3. Style it​

Theme attributes are forwarded to the chat document, which validates every one of them. A value it doesn't recognise falls back to the default rather than breaking the widget.

AttributeDefaultEffect
data-accent#2563ebSend button and the visitor's own messages. Text on it is switched between light and dark automatically.
data-bg#ffffffPage background. A dark value flips the derived text and border colours.
data-fgderived from data-bgText colour
data-surfacederived from data-bgBackground of the agent's messages
data-radius12Corner radius in px, 0–32
data-fontsystem stackCSS font-family list
data-chromeonoff hides the header
<script src="https://agent-api.verinfast.com/v1/share/embed.js"
data-token="dsl_…"
data-accent="#0f766e"
data-bg="#0b1220"
data-radius="4"
data-font="Inter, system-ui, sans-serif"
data-title="Ask Verinfast"></script>

Colours are hex only (#rgb or #rrggbb) — no rgb(), no named colours, no CSS variables. Fonts must already be available on the visitor's machine. The widget is a separate document and loads no external resources at all, so a webfont your page imports is not available inside it; name it first in the stack and give it a system fallback.

If you need more control than this, skip the widget and build your own UI.

Markdown in the agent's replies​

Agents write markdown, and the widget renders it: headings, bold, italic, strikethrough, inline code and fenced code blocks, ordered and unordered lists (nested included), blockquotes, horizontal rules, tables, and links. Everything inherits your theme attributes — code blocks and table headers use a tint of the surrounding colour, so they work on a light or a dark data-bg without extra configuration. A wide table or a long code line scrolls inside its own message rather than stretching the frame.

Two deliberate exceptions:

  • Links open in a new tab (rel="noopener noreferrer"), and any target that isn't http:, https:, mailto: or tel: is dropped — the link text is still shown, just not clickable.
  • Images render as a link to the image, labelled with its alt text. The chat document's content security policy allows no remote resources at all, so an <img> pointing off-origin could only ever be a broken icon.

Raw HTML in a reply is shown as text, never as markup. What the visitor types is always displayed verbatim — their asterisks stay asterisks.

The widget's wording​

Four pieces of visible text can come either from the agent's record or from the script tag.

AttributeFalls back toLimit
data-titlethe agent's name80 chars
data-subtitlethe agent's description140 chars
data-placeholderMessage <agent name>…80 chars
data-greetingthe agent's Initial Greeting, or Say hello to <agent name> to get started.400 chars

The attribute wins whenever it is present. Each is resolved in the same order — script tag first, then the agent record, then a built-in default — so an attribute on the page makes the corresponding field in Agent → Settings unreachable for that embed. This is the usual explanation for editing a field in Settings and seeing no change on the site; see Troubleshooting.

Which layer to use is a real choice, not a formality:

  • Leave the attributes off to manage all four from Agent Settings. One agent, one voice, changeable without a site deploy. Best when the same widget appears in several places.
  • Set them on the page when the wording should differ per page — a pricing page and a docs page seeding different questions off one share link. The cost is that the copy now lives in your site's repo and will drift from the agent record unless someone maintains both.

Values are whitespace-collapsed and silently truncated at the limits above, so write to the limit rather than past it. An empty attribute (data-greeting="") is treated as absent and falls back — it does not blank the text.

The greeting is shown, not sent. In the widget it is placeholder text for the empty chat: it is not sent as a message, and it is not part of the system prompt the agent is given, so it costs no tokens and does not steer the reply. It disappears as soon as the first message is sent. Wording meant to shape how the agent replies belongs in the agent's description prompt.

4. Drive it from your page​

DelegateAgent.open(); // launcher mode
DelegateAgent.close();
DelegateAgent.toggle();
DelegateAgent.focus(); // inline mode — focus the composer
DelegateAgent.destroy();

The widget posts messages to the host page. Check the origin before trusting one:

window.addEventListener("message", (e) => {
if (e.origin !== "https://agent-api.verinfast.com") return;
if (e.data?.source !== "delegate-agent") return;
// e.data.type is "ready", "close", or "idle"
});

ready fires when the chat has loaded, close when the visitor dismisses the panel, and idle when a reply has finished streaming — useful for a "new message" badge when the panel is shut.

Restricting who can embed​

allowed_origins is served as a frame-ancestors directive on the embed document, so the browser refuses to render your agent inside a page you haven't named. Wildcards work at the subdomain level:

{ "allowed_origins": ["https://verinfast.com", "https://*.verinfast.com"] }

Origins are normalised when stored: lower-cased, default ports dropped, duplicates removed. A scheme and host are required; paths are rejected.

An empty allow-list means any site may embed the link. Set one before the token goes into a public page.

You can change the list without minting a new token — the old one is already pasted into your HTML:

PATCH /v1/agent/{agent_id}/share-links/{link_id}
{ "allowed_origins": ["https://verinfast.com"] }

Sending [] clears the restriction.

What this is and isn't. frame-ancestors is enforced by the visitor's browser, which makes it a real control against your agent being framed on someone else's site — and no control at all against a script that skips the browser entirely and calls the API directly. The controls that hold there are the per-link daily cap, your agent's budget caps, and revocation.

What a visitor can do​

Share-link sessions run with the read-only toolset. The agent cannot change its own configuration, skills, routines, memory, or identity, and tenant-level custom HTTP tools are dropped — their credentials belong to your tenant, not to a stranger on the internet.

It can still read and write workspace files and run code in its sandbox.

One workspace, many visitors

Sessions are isolated from each other: a visitor can only ever load a session started by the same link, and cannot list or read anyone else's. The agent's workspace is not isolated — it is shared across every session that agent runs. A file one visitor has the agent write can be read by the next.

For a public-facing agent, use an agent dedicated to that purpose, and don't put anything in its workspace you wouldn't publish.

File uploads​

Visitors cannot attach files unless the link says they can. Turn it on per link in Agent → Settings → Share Links (the Uploads column toggles an existing link without re-minting its token), or at mint time:

POST /v1/agent/{agent_id}/share-links
{ "label": "Support widget", "allow_uploads": true }

With uploads on, the composer grows a paperclip button. An attached file is uploaded into the agent's workspace under /attachments/<session>/, exactly like an attachment from a signed-in user, and the agent is told where to find it. The caps are much tighter than the signed-in ones — 3 files per message, 10MB each, 20MB in total — because the sender is anonymous.

Uploads land in the shared workspace

The workspace warning above applies with more force here: a file one visitor attaches sits in the same workspace the next visitor's session can read. Leave uploads off unless the agent needs them, and prefer a dedicated agent when it does.

Limits, errors, and lifetime​

ConditionWhat the visitor sees
Link revoked, expired, or unknownThe chat renders "no longer available" (HTTP 404)
Daily message cap reached (200 messages per link per UTC day by default)"Reached its message limit for today" (HTTP 429)
Agent or tenant budget cap reachedAn error message in the chat
A message longer than 20,000 charactersRejected before it reaches the agent
An attachment sent to a link with uploads off"File uploads are not enabled for this link" (HTTP 403)
An attachment over the per-file, count, or total capThe limit that was hit, in the chat (HTTP 400)

Revoking a link takes effect on the next request — there is no cache to wait out. Sessions started by that link stop working with it.

Conversations survive a page reload via sessionStorage, keyed to the link. That storage is partitioned or blocked outright for third-party frames in some browsers, so treat scrollback as a nicety: when it isn't available, the visitor simply starts fresh. Nothing else depends on it.

Configuring the agent behind the widget​

A few agent settings behave differently once the people talking to it are anonymous strangers rather than signed-in colleagues.

Set an initial greeting​

Agent → Settings → General → Initial Greeting is what an embed shows when no data-greeting is on the page. Leave it blank and visitors get Say hello to <agent name> to get started. — which names an internal label and tells a first-time visitor nothing about what the agent knows.

The empty chat is also the only place you get to seed questions, and visitors who don't know what to ask usually don't ask. Name the two or three things the agent is actually good at.

Leave bootstrap instructions empty​

Bootstrap and public agents don't mix

Bootstrap instructions fire exactly once in an agent's lifetime, on its first session — whoever that turns out to be. For an agent behind a public widget, that is very likely an anonymous visitor rather than you.

Worse, it cannot do anything useful there: share-link sessions run read-only, so update_agent_settings is stripped from the toolset and the agent cannot save whatever the bootstrap conversation establishes. The instructions are spent, the result is discarded, and one visitor gets a setup interview instead of an answer.

Put standing instructions in the Description Prompt instead. It is injected into every session, which is what "configure this agent" almost always means.

Identity onboarding is handled for you​

An agent whose user and company identity are unset normally opens a conversation by introducing itself and asking who it is talking to. That does not happen in a share-link session: an anonymous visitor has no identity record to write to, and manage_identity is not in the read-only toolset, so the prompt is suppressed rather than asked with nowhere to put the answer.

There is nothing to configure. It is worth knowing the rule, though, because the same agent opened from the Delegate UI by a signed-in user will still ask — identity is stored per user per organisation, so the same person is asked once in each organisation they use the agent from.

Troubleshooting​

I changed the greeting, title, or subtitle in Settings and the widget still shows the old text. A data-* attribute on the script tag overrides the agent record — check the page source for data-greeting, data-title, data-subtitle, and data-placeholder, and remove the one you meant to control from Settings. To see what the agent record itself holds, ask the token:

curl -s "https://agent-api.verinfast.com/v1/share/<token>" | jq
{
"agent_id": "…",
"name": "Support",
"description": "…",
"initial_greeting": "Ask me about pricing, SSO, or data residency."
}

If initial_greeting is your new text, the record is fine and the override is on your page. If it still shows the old text, check agent_id: a token resolves to exactly one agent, so a link minted against a different agent — or another organisation's copy of it — ignores edits made to the agent you were looking at. Compare that agent_id with the agent whose Settings you changed, and the data-token on the page with the link under its Share Links.

Do I need to wait for a cache? For agent-record changes, no — the chat document is served no-store and rebuilt per load. The embed.js script is cached for five minutes, so only changes to the widget itself take that long.

The wording is cut off. Each text attribute is truncated at the limit in The widget's wording. Truncation is silent.

The chat renders "no longer available". The link is revoked, expired, or the token is wrong. This is a real HTTP 404 from the embed document, and it looks different from an allow-list rejection: a site missing from allowed_origins is blocked by the browser before anything renders, leaving an empty frame and a frame-ancestors violation in the console.

Build your own UI​

The widget is a convenience over a documented HTTP API. If you want your own interface, here it is.

Send a message​

POST /v1/share/{token}/perform_task
Content-Type: application/json

{ "task": "Do you support SSO?", "session_id": null }

Omit session_id (or send null) to start a conversation; pass the id from a previous reply to continue one. The last 20 messages of that session are replayed to the agent as history.

Read the stream​

The response is newline-delimited JSON, not SSE — one object per line, each shaped {"type": …, "data": {…}, "timestamp": …}. Read it with a streaming reader and split on \n; skip blank lines, and skip a line that fails to parse rather than aborting the conversation.

typeWhat it means
task_startedThe run has begun. Emitted before the first model call — the earliest thing you can render.
planningThe agent is deciding what to do
tool_callA tool is being invoked; data.name is the tool
tool_resultThat tool returned
thinkingReasoning output, when the model produces it
messageAssistant text in data.content
keepaliveHeartbeat, roughly every 30s, so proxies hold the connection. Ignore it.
message_savedThe assistant message has been persisted
throttledThe run is being slowed by a budget guardrail
task_completedTerminal. data.session_id is the id to send back next time.
task_cancelledTerminal
errorTerminal. data.error carries the message.

Treat unknown types as pass-through — the list grows, and a client that throws on an unrecognised type breaks on our deploys, not yours.

data.content on a message event is markdown — the same text the widget renders. Render it with the markdown library you already use, and sanitise the result; the text comes from a model, so treat it as untrusted.

message events carry a whole block of text, not token-by-token deltas. There is no incremental token streaming today, so a typewriter effect isn't available; show a thinking state (and tool_call activity, which is genuinely informative) until the first message arrives.

Other endpoints​

GET /v1/share/{token} → agent name, description, greeting
GET /v1/share/{token}/sessions/{session_id} → that session's messages
GET /v1/share/{token}/stream_session/{session_id} → re-attach to a run in progress

If a stream drops mid-answer the run keeps going server-side. Re-attach with stream_session while it's still running (404 once it isn't), or fetch the session afterwards to collect the finished reply. Either way the session survives — the visitor never has to re-ask.

One constraint​

These endpoints are on Delegate's origin, so a browser calling them from your site is making a cross-origin request, and your origin has to be on the API's CORS allow-list. That list is managed by us, per deployment — get in touch if you need one added. Calls from your server have no such constraint, and neither does the widget, which runs on our origin to begin with.