开发 dsh-plugin-deepseek-browser-use:会话、作业、文件与图片
上一篇已经把工具请求送到了 Java 服务。真实使用还有几个问题:两个会话会不会操作同一个任务?同一会话同时发两个命令怎么办?HTTP 已返回,但异步批次还没结束怎么办?本篇结合插件 0.2.0 的实际实现,逐步补齐这些核心能力。
1. 为什么不能只保存一个全局 taskId
假设 A 会话打开订单页面,B 会话打开搜索页面。如果两个会话共用一个任务,A 下一次点击可能落在 B 的页面上。把 task ID 交给模型填写同样不合适:它既可能填错,也可能引用另一个会话的任务。
当前实现用 Map<Owner, Entry> 保存状态,Owner 是通过工具入口验证过的活跃 Agent 对象。以下是 src/sessions.ts 完整文件;先通读整体,再按后续小节理解各个分支:
import { randomInt } from 'node:crypto';
import { BrowserClient, requireSuccess, type Envelope, type Params } from './client.js';
export interface Owner { ctx: { effect(callback: () => () => Promise<void>, label?: string): () => Promise<void> }; }
interface Entry { id: number; tail: Promise<unknown>; started: boolean; closed: boolean; abort: AbortController; cleanup: () => Promise<void>; jobs: Set<string>; activeJob?: string; batchUnknown?: boolean; }
export interface SessionOptions { browser: string; headless: boolean; }
/** Live owner identity, not a model-supplied task ID, controls resource access. */
export class BrowserSessions {
private entries = new Map<Owner, Entry>();
private disposed = false;
constructor(readonly client: BrowserClient, readonly options: SessionOptions) {}
private entry(owner: Owner): Entry {
if (this.disposed) throw new Error('Browser plugin is disposed');
const previous = this.entries.get(owner);
if (previous) return previous;
// 48 random bits: exact JSON/JavaScript integer, unlike server Snowflake IDs.
const entry: Entry = { id: randomInt(1, 2 ** 48 - 1), tail: Promise.resolve(), started: false, closed: false,
abort: new AbortController(), cleanup: async () => {}, jobs: new Set() };
this.entries.set(owner, entry);
entry.cleanup = owner.ctx.effect(() => async () => {
entry.closed = true;
entry.abort.abort(new Error('Browser session disposed'));
await entry.tail.catch(() => {});
if (entry.batchUnknown) throw new Error(`Task ${entry.id}: batch submission outcome unknown. Inspect service list_jobs and finish/cancel the batch before manually closing this task. Cleanup did not close potentially active work.`);
if (entry.started) {
const signal = AbortSignal.timeout(this.client.options.timeoutMs);
if (entry.activeJob) {
requireSuccess(await this.client.command(entry.id, 'cancel_job', { jobId: entry.activeJob }, signal));
for (;;) {
const status = requireSuccess(await this.client.command(entry.id, 'get_job', { jobId: entry.activeJob, includeResult: false }, signal));
if ((status.data as Params)?.status !== 'running') break;
await new Promise(resolve => setTimeout(resolve, 100));
signal.throwIfAborted();
}
}
requireSuccess(await this.client.command(entry.id, 'close', {}, signal));
}
this.entries.delete(owner);
}, 'deepseek-browser-use.session');
return entry;
}
run(owner: Owner, method: string, params: Params, signal: AbortSignal): Promise<Envelope> {
signal.throwIfAborted();
const entry = this.entry(owner);
const combined = AbortSignal.any([signal, entry.abort.signal]);
const operation = entry.tail.then(async (): Promise<Envelope> => {
combined.throwIfAborted();
if (entry.closed) throw new Error('Browser session is closed');
if (entry.batchUnknown) throw new Error(`Task ${entry.id}: batch submission outcome unknown; inspect service list_jobs before manually recovering. This Session will not submit more browser actions.`);
if (method === 'get_job' || method === 'cancel_job') {
if (typeof params.jobId !== 'string' || !entry.jobs.has(params.jobId)) throw new Error('Job does not belong to this Session');
} else if (entry.activeJob) {
throw new Error(`Browser batch ${entry.activeJob} is pending. Use dsb_job until it finishes before issuing other commands.`);
}
if (method === 'close' && !entry.started) return { ok: true, data: { id: entry.id, closed: true } };
if (!entry.started) {
// Remember ownership before HTTP, so timeout during start is still cleaned up.
entry.started = true;
const started = await this.client.command(entry.id, 'start', { ...this.options }, combined);
if (!started.ok) entry.started = false;
requireSuccess(started);
if (method === 'start') return started;
} else if (method === 'start') {
return { ok: true, data: { id: entry.id, reused: true } };
}
let result: Envelope;
try { result = await this.client.command(entry.id, method, params, combined); }
catch (error) {
if (method === 'commands') entry.batchUnknown = true;
throw error;
}
const data = result.data as Params | undefined;
if (method === 'commands' && result.ok) {
if (typeof data?.jobId !== 'string') {
entry.batchUnknown = true;
throw new Error(`Task ${entry.id}: async batch response missing jobId; inspect service before resubmitting`);
}
entry.jobs.add(data.jobId);
entry.activeJob = data.jobId;
}
if (method === 'get_job' && result.ok && data && ['done', 'failed', 'cancelled'].includes(String(data.status)) && params.jobId === entry.activeJob) {
entry.activeJob = undefined;
}
if (method === 'close' && result.ok) entry.started = false;
return result;
});
entry.tail = operation.catch(() => {});
return operation;
}
async dispose(): Promise<void> {
this.disposed = true;
const results = await Promise.allSettled([...this.entries.values()].map(entry => entry.cleanup()));
const errors = results.filter(r => r.status === 'rejected').map(r => r.reason);
if (errors.length) throw new AggregateError(errors, 'Browser session cleanup failed');
}
}
这份文件可以在原工程中编译;它依赖上一篇介绍的 client.ts。Owner 接口只声明会话管理需要的 ctx.effect,不把整个 Harness Agent 类型耦合到 HTTP 会话管理器里,便于独立测试。
2. 先读懂 Entry 中的字段
| 字段 | 含义 | 缺少它会怎样 |
|---|---|---|
id | 发送到 Java 的任务 ID | 无法把后续命令关联到自己的页面 |
tail | 当前队列尾部的 Promise | 同会话并发动作可能乱序 |
started | 任务是否已启动或可能已启动 | 重复 start,或漏掉需要清理的任务 |
closed | Agent 作用域是否已关闭 | 排队调用可能在释放之后继续操作 |
abort | 本会话的取消控制器 | 释放作用域时无法通知等待中的请求 |
cleanup | 注册到作用域的清理函数 | 插件卸载时无法统一释放资源 |
jobs | 本会话创建过的 job ID 集合 | 可能查询或取消别人的作业 |
activeJob | 尚未观察到终态的作业 | 普通命令可能插入仍运行的批次 |
batchUnknown | 批次提交结果无法确认 | 响应丢失后可能重复提交整个批次 |
randomInt(1, 2 ** 48 - 1) 在 48 位范围内选择正整数,可被 JavaScript number 精确表示,也可传入 Java long。这里避免直接使用可能超过 JavaScript 安全整数范围的服务端 Snowflake ID;随机值不是绝对无碰撞保证,也不是认证凭据。
任务归属隔离不等于登录态隔离:多个任务仍可能使用同一个浏览器 profile 的 Cookie。需要不同账号互不干扰时,应使用独立服务和 profile。
3. 顺序队列为什么写成两个 Promise
关注 run 中这三行的关系:
const operation = entry.tail.then(async (): Promise<Envelope> => {
// 在这里检查取消、启动任务、发送本次命令。
});
entry.tail = operation.catch(() => {});
return operation;
这是结构示意,注释位置在完整文件中有实际实现,不是可直接替换的方法。
新的 operation 接在旧 tail 后面,因此同一 Owner 的两个请求按加入队列的顺序执行。返回给调用者的是原始 operation,调用者仍能看到失败。保存给下一项的 tail 则通过 catch 恢复为可继续等待的 Promise,避免一次失败让后续所有调用都跳过执行。
排队前检查一次取消,进入队列后还要检查一次。假设第二个点击在等待第一个请求时被取消,只在函数入口检查就会漏掉这个变化,最终仍把取消的点击发给浏览器。
4. 懒启动与 close 的两种含义
首次调用 dsb_state 或 dsb_navigate 时,started 为 false,管理器先发送 start,成功后再发送目标命令。后续 dsb_start 返回 reused:true,不会重复创建任务。
为什么在 start 请求之前先设置 started = true?因为 HTTP 超时不代表服务端没有创建任务。提前记下“可能已启动”,作用域清理时才能尝试关闭它;若收到明确的 ok:false,则恢复为 false。
显式调用 dsb_close 只把 started 改回 false,保留活跃 Agent 的 Entry,所以下次可以重新 start。注意:若还存在未观察到终态的批次,dsb_close 会被直接拒绝(activeJob 检查在 close 分支之前),需要先用 dsb_job 观察到终态。Agent 作用域释放时才设置 closed = true,并取消信号、等待队列和清理任务。这两个动作语义不同。
5. 通用命令为什么还需要允许列表
把模型传来的任意 method 转给服务很方便,但会绕过任务生命周期和异步作业管理。当前工具层先检查准确命令名,再排除必须走专用工具或不公开的操作:
const forbidden = new Set(['start', 'close', 'shutdown', 'cleanup', 'get_config', 'list_tasks', 'list_jobs', 'get_job', 'cancel_job', 'run_recipe', 'commands']);
export function validateCommand(method: string): void {
if (!commandNames.includes(method) || forbidden.has(method)) {
throw new Error(`Command ${method} is not an exposed task command. Use dedicated lifecycle/job tools; service-wide commands and recipes are not exposed.`);
}
}
commandNames 来自 src/commands.ts,是 Java CommandTable.java 的名称快照。它不是每次联网动态发现的服务能力。新增服务命令时要更新快照,并运行现有一致性测试。
这里禁止通用入口直接调用 start/close/get_job/cancel_job,是为了统一管理会话;禁止 commands/run_recipe,是为了避免嵌套动作绕过逐项验证;禁止 shutdown/cleanup 等全局操作,是为了避免影响其他会话。这些限制不是页面 JavaScript 的安全沙箱。
6. 异步批次:队列结束不等于浏览器工作结束
以下是 src/tools.ts 中通用命令、批次和作业工具的实际注册代码,依赖已有 add、command 和 object:
add('command', 'Execute an exact supported task-scoped command. Use dsb_methods and the skill for names/parameters. No lifecycle, global management, nested batches or recipes.',
z.strictObject({ method: z.string().min(1), params: object.optional() }), (a, e) => { validateCommand(a.method); return command(e, a.method, a.params ?? {}); });
add('batch', 'Submit sequential task commands as a background job. Poll dsb_job before issuing other actions. Do not batch indices whose validity depends on earlier page changes.', z.strictObject({
commands: z.array(z.strictObject({ method: z.string(), params: object.optional() })).min(1).max(200),
stopOnError: z.boolean().optional(),
}), (a, e) => {
for (const step of a.commands) validateCommand(step.method);
return command(e, 'commands', { commands: a.commands.map(step => ({ [step.method]: step.params ?? {} })), async: true, stopOnError: a.stopOnError ?? true });
});
add('job', 'Read or request cooperative cancellation of a job created by this Session. Cancellation does not undo an already delivered action.',
z.strictObject({ jobId: z.string().min(1), cancel: z.boolean().optional() }), (a, e) => command(e, a.cancel ? 'cancel_job' : 'get_job', { jobId: a.jobId, ...(a.cancel ? {} : { includeResult: true }) }));
模型更容易填写这种结构:
{
"commands": [
{ "method": "get_title" },
{ "method": "get_url" }
],
"stopOnError": true
}
服务需要的 params 则是:
{
"commands": [{ "get_title": {} }, { "get_url": {} }],
"async": true,
"stopOnError": true
}
map(step => ({ [step.method]: step.params ?? {} })) 完成两种格式的转换;方括号表示用变量值作为属性名。批次中的每一项先经过 validateCommand,所以不能把被禁止的命令藏进数组。
服务返回 job ID 后,这次 HTTP 就结束了,但批次可能仍在执行。BrowserSessions 因此保存 activeJob,暂时拒绝其他普通动作,只允许查询或取消本会话拥有的作业。查询到 done/failed/cancelled 才放行。
还有一条同样会卡死的路径要记住:作业结果只保存在服务端内存里(只保留最近 50 个,服务重启即丢)。作业查不到时 get_job 返回 ok:false,而实现只在 ok:true 时才清除 activeJob,于是这个会话的普通浏览器动作会被一直拒绝,错误文案却还在提示继续用 dsb_job。遇到“没有这个任务”时不要反复轮询,按 batchUnknown 的方式收尾(交给运维核查服务端作业,然后新建 Harness 会话)。
| 观察到的情况 | 会话处理 | 使用者下一步 |
|---|---|---|
| 提交成功,有 job ID | 保存归属并锁住普通动作 | 调用 dsb_job |
| 查询仍在运行 | 保留 activeJob | 稍后继续查询 |
| 查询到终态 | 清除 activeJob | 检查内部结果,再继续操作 |
| 取消请求成功 | 尚不把它当作执行终态 | 继续查询确认停止 |
| 提交响应丢失或成功响应没有 job ID | 设置 batchUnknown | 由运维核查服务端作业,不能盲目重放 |
作业查询外层 ok:true 只说明“查到了”,执行是否成功还要看 data.status 和内部 data.ok。页面索引会随着前一步操作变化,因此不要把依赖新索引的多次点击盲目塞进同一个批次。
7. 资源释放为什么不能直接 shutdown
完整 entry 方法中注册的清理逻辑依次做:标记关闭、发出取消、等待本会话队列、处理未结束的作业、关闭自己的 task、删除 Map 条目。
有活跃作业时先 cancel_job,再查询到不再 running,才 close。取消是协作式的,不会撤销已点击的按钮。清理等待受 timeoutMs 限制,不能永远阻塞卸载。
batchUnknown 时直接报告清理失败,保留 task ID,避免关闭仍可能运行的作业。运维需要在 Java 服务端 list_jobs 中匹配 browserId,等待或取消作业后关闭任务,然后使用新的 Harness 会话。插件通用工具本身不暴露全局 list_jobs。
插件级 dispose() 使用 Promise.allSettled,一个会话清理失败不会阻止尝试清理其他会话,最后用 AggregateError 汇总失败。整个路径不调用共享服务的 shutdown。
8. 上传文件:传字节而不是把本机路径交给远程服务
这是 src/tools.ts 中的上传工具片段:
add('upload', 'Upload a file from the Harness host to a file input. localPath is on the Harness host; bytes are transferred over HTTP even if browser service is remote.',
z.strictObject({ ...target, localPath: z.string().min(1) }), async (a, e) => {
checkTarget(a);
const info = await stat(a.localPath);
if (!info.isFile() || info.size > options.maxUploadBytes) throw new Error('Upload must be a regular file within maxUploadBytes');
const bytes = await readFile(a.localPath, { signal: e.signal });
if (bytes.length > options.maxUploadBytes) throw new Error('Upload exceeds maxUploadBytes');
const { localPath, ...locator } = a;
return command(e, 'upload_file', { ...locator, filename: basename(localPath), contentBase64: bytes.toString('base64') });
});
用到的 Node import 是:
import { readFile, stat } from 'node:fs/promises';
import { basename } from 'node:path';
localPath 属于 Harness 主机,Java 服务不一定在同一台机器,所以请求发送 filename/contentBase64。stat 先检查普通文件及大小,读完再检查实际字节数,捕获检查与读取之间文件变大的情况;当前实现仍是内存读取,不是严格有界的流式上传。base64 还会增加约三分之一的传输体积。
const { localPath, ...locator } = a 从参数中分离出本机路径,避免把它误当作页面定位参数传给服务。index/selector 的二选一检查继续复用基础工具逻辑。
9. 截图:图片路径和模型看到图片是两件事
这是截图工具的实际代码:
add('screenshot', 'Capture the page or a selector. Set view=true to attach image for a verified vision-capable model; otherwise return a saved screenshot URL.',
z.strictObject({ fullPage: z.boolean().optional(), selector: z.string().optional(), view: z.boolean().optional() }), async (a, e) => {
const attachments = ctx.get('attachments');
if (a.view) {
if (!attachments) throw new Error('view=true requires the Harness attachment service');
const route = e.agent?.session.requestHeader()?.config;
const provider = route?.provider ?? e.agent?.options.provider;
const model = route?.model ?? e.agent?.options.model;
const llm = ctx.get('llm');
if (!provider || !model || !llm) throw new Error('Cannot verify the active model image capability; use view=false');
const info = await llm.resolveModelInfo(provider, model, e.signal);
if (!info.inputModalities?.includes('image')) throw new Error('Current model does not declare image input; use view=false and DOM text tools');
}
const { view, ...params } = a;
const result = await command(e, 'screenshot', { ...params, inline: view ?? false });
if (!result.ok || !view) return result;
const data = result.data as Params;
if (typeof data?.base64 !== 'string') throw new Error('Screenshot response missing base64');
e.signal.throwIfAborted();
const attachment = await attachments!.saveImage({ data: Buffer.from(data.base64, 'base64'), mediaType: 'image/png', name: 'browser-screenshot.png' });
const { base64: _bytes, ...metadata } = data;
return { ...result, data: metadata, attachment: attachment as unknown as Params };
});
默认 view:false 只返回截图元信息、路径或 URL,减少图片传输。view:true 时流程是:
- 读取当前调用路由的 provider/model,缺失时使用 Agent 配置;确认宿主附件服务和 LLM 服务可用。
- 通过模型信息检查
inputModalities是否含image,不凭模型名字猜测视觉能力。 - 请求
screenshot(inline:true),拿到 base64 后解码为字节。 - 用
attachments.saveImage保存为宿主附件,移除规范结果里的 base64。 - 由基础篇的
output.render生成 image 内容块。
如果直接把 base64 放进文本,模型不会因此得到正常的图片输入,还会占用大量上下文。没有图片能力时,可以继续使用默认截图留档和 DOM 文本工具。
10. 用测试证明这些设计有效
工程的 test/plugin.test.mjs 已包含以下测试,运行 npm test 可验证:
| 测试情形 | 要证明的行为 |
|---|---|
| 同一 Owner 并发 execute_js | 实际 HTTP 不重叠,且只 start 一次 |
| 两个不同 Owner | 使用不同的安全整数 task ID |
| 排队点击被取消 | 没发出点击,下一条请求仍正常执行 |
| 批次运行时再点击 | 拒绝插入动作,观察终态后恢复 |
| 另一个 Owner 查询 job | 拒绝访问非本会话作业 |
| 批次响应被截断 | 冻结会话,不重发,不假装清理成功 |
| 上传中文文件 | 传的是实际文件字节,文件名保留 |
| 截图展示 | 检查模型能力,保存附件,不把 base64 留在结果中 |
这些测试大多使用假 HTTP 服务,验证的是协议和状态管理。真实浏览器、目标 Harness 的激活、Web UI 图片显示是不同层次的验证,不能互相代替。
上一篇:HTTP 客户端与基础工具。下一篇:构建并安装到 DeepSeek Harness。
