> ## Documentation Index
> Fetch the complete documentation index at: https://docs.xpertai.cn/llms.txt
> Use this file to discover all available pages before exploring further.

# 插件权限设计指南

<Tip>
  在 meta 中设计一套可以向核心系统反向请求 大模型 调用权限的（如请求一个 OpenAI Client实例）的配置，例如在 meta中声明大模型的权限请求，在 configSchema 中设计属性让用户配置具体的模型提供商和模型名称，核心系统在识别到权限请求时会根据 config 中用户配置的具体参数实例化大模型 Client 通过 config 传入到 understandImages 这个方法：
</Tip>

👌明白了，你希望 **在插件 Strategy 的 meta 中声明依赖的大模型权限请求**，核心系统解析后根据用户配置（模型提供商 + 模型名称 + 其他参数）自动注入一个 `client` 到 `config`，这样插件不用自己创建 OpenAI / Anthropic / 其他 VLM 客户端。

我帮你重新设计 `OpenAIGPT4VStrategy` 的 meta 和 configSchema，使它能 **反向请求核心系统提供的 Client**。

***

## 新版 Strategy

```ts theme={null}
import { Injectable } from '@nestjs/common';
import { Document } from 'langchain/document';
import {
  ImageUnderstandingStrategy,
  IImageUnderstandingStrategy,
  TImageUnderstandingFile,
  TImageUnderstandingResult
} from './strategy.interface';

@Injectable()
@ImageUnderstandingStrategy('openai-gpt4v')
export class OpenAIGPT4VStrategy implements IImageUnderstandingStrategy<any> {
  readonly meta = {
    name: 'openai-gpt4v',
    label: { en: 'OpenAI GPT-4 Vision', zh: 'OpenAI GPT-4 图像理解' },

    /**
     * 权限声明：告诉核心系统，这个插件需要一个 "llm" 权限，
     * 提供方是 openai，能力是 vision。
     * 核心系统识别后会根据 config.provider + config.model 实例化对应的 Client，
     * 并注入到 config.client
     */
    permissions: [
      {
        type: 'llm',
        provider: 'openai',
        capability: 'vision',
      },
    ],

    /**
     * 配置 schema，用户需要填写 provider 和 model，
     * 核心系统会根据 provider + model + 用户账号 API Key，
     * 实例化好 Client 注入。
     */
    configSchema: {
      type: 'object',
      properties: {
        provider: { 
          type: 'string', 
          enum: ['openai'], 
          description: '模型提供商' 
        },
        model: { 
          type: 'string', 
          default: 'gpt-4-vision-preview', 
          description: '大模型名称' 
        },
        prompt: { 
          type: 'string', 
          default: 'Describe this image in detail.', 
          description: '提示词（可以自定义理解任务）' 
        }
      },
      required: ['provider', 'model']
    },

    icon: {
      svg: '<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 20 20"><circle cx="10" cy="10" r="8" stroke="black" fill="none"/><text x="6" y="14">VLM</text></svg>',
      color: '#10a37f'
    }
  };

  async validateConfig(config: any): Promise<void> {
    if (!config.client) {
      throw new Error('Missing injected LLM client in config. Core system must inject client.');
    }
  }

  async understandImages(
    files: TImageUnderstandingFile[],
    config: any
  ): Promise<TImageUnderstandingResult[]> {
    const client = config.client; // ✅ 已由核心系统注入
    const results: TImageUnderstandingResult[] = [];

    for (const file of files) {
      const description = await this.runGPT4V(client, file.path, config);

      const doc = new Document({
        pageContent: description,
        metadata: {
          chunkId: `img-${Date.now()}-${Math.random().toString(36).slice(2, 8)}`,
          parentChunkId: file.parentChunkId,
          imagePath: file.path,
          source: file.filename,
          type: 'vlm',
          engine: config.model
        }
      });

      results.push({
        docs: [doc],
        metadata: { engine: config.model, file: file.filename }
      });
    }

    return results;
  }

  private async runGPT4V(client: any, imagePath: string, config: any): Promise<string> {
    // client 已经是 openai client，由核心系统实例化
    const response = await client.chat.completions.create({
      model: config.model,
      messages: [
        {
          role: 'user',
          content: [
            { type: 'text', text: config.prompt || 'Describe this image in detail.' },
            { type: 'image_url', image_url: { url: imagePath } }
          ]
        }
      ],
      max_tokens: 500
    });

    return response.choices[0]?.message?.content || '';
  }
}
```

***

## 🔑 核心系统如何处理

* 解析插件的 `meta.permissions`，发现需要：

  ```json theme={null}
  { "type": "llm", "provider": "openai", "capability": "vision" }
  ```
* 根据用户 `config.provider = "openai"`, `config.model = "gpt-4-vision-preview"`
  实例化 `OpenAI({ apiKey: userApiKey })`
* 将 client 注入到 `config.client`
* 调用 `strategy.understandImages(files, config)`

***

## 📌 使用示例

```ts theme={null}
const strategy = this.imageUnderstandingRegistry.get('openai-gpt4v');
const result = await strategy.understandImages(
  [
    { 
      path: 'https://example.com/chart.png', 
      filename: 'chart.png', 
      extname: 'png', 
      parentChunkId: 'uuid-123' 
    }
  ],
  { 
    provider: 'openai',
    model: 'gpt-4-vision-preview',
    prompt: '请用中文详细描述这张图片。',
    client: openaiClient // ✅ 核心系统注入
  }
);

console.log(result[0].docs[0].pageContent);
```

***

这样插件本身不依赖任何 `openai` 包装逻辑，**只声明权限 + 要求 provider/model**，核心系统就能注入合适的 client。

要不要我帮你写一个 **抽象的权限声明接口 (例如 ILLMPermissionRequest)**，这样 OCR/VLM/Embedding 插件都可以统一声明依赖？
