Skip to content

GPT Image 生图与编辑 ​

生成图片 ​

以下示例直接从 ~/.codex/auth.json 读取 OPENAI_API_KEY,并将它作为 WAIHUB_API_KEY 使用。不要检查 WAIHUB_API_KEY 是否存在,也不要在终端输出、日志或代码仓库中暴露密钥。

bash
export WAIHUB_API_KEY="$(jq -r '.OPENAI_API_KEY' ~/.codex/auth.json)"

curl --fail-with-body --location 'https://waihub.top/v1/images/generations' \
  --header "Authorization: Bearer ${WAIHUB_API_KEY}" \
  --header 'Content-Type: application/json' \
  --data '{
    "model": "gpt-image-2.5",
    "prompt": "画一只可爱的猫抱着水獭",
    "size": "1024x1024",
    "quality": "xhigh",
    "output_format": "png",
    "n": 1
  }' \
  -o image_response.json

# macOS 使用 -D;Linux 使用 -d
jq -e -r '.data[0].b64_json' image_response.json | base64 -D > cat_otter.png

编辑图片 ​

编辑接口使用 multipart/form-data,输入图片字段是 image:

bash
export WAIHUB_API_KEY="$(jq -r '.OPENAI_API_KEY' ~/.codex/auth.json)"

curl --fail-with-body --location 'https://waihub.top/v1/images/edits' \
  --header "Authorization: Bearer ${WAIHUB_API_KEY}" \
  --form 'model=gpt-image-2.5' \
  --form 'prompt=把背景改成柔和的浅蓝色摄影棚背景,保持主体不变' \
  --form 'image=@input.png;type=image/png' \
  --form 'output_format=png' \
  -o edit_response.json

jq -e -r '.data[0].b64_json' edit_response.json | base64 -D > edited.png

让 AI 自动创建生图 Skill

把下面整段指令复制给本地的 Codex、Claude Code 或其他支持 Skill 的 AI 编程助手,它就可以按照本页说明为你完成配置:

text
请参考 https://docs.waihub.top/guide/extension/gpt-image-gen 页面中的最新接口说明,在我的本地环境中创建一个名为 gpt-image-gen 的生图与图片编辑 Skill。

要求:
1. 先查看本页最前面的“生成图片”和“编辑图片”curl 调用示例,并严格按其中的请求格式、鉴权方式和 Base64 解码方式实现。使用 Waihub 的 OpenAI 兼容接口和 gpt-image-2.5 模型(默认);也支持 gpt-image-2.5-flare 与 gpt-image-2.5-sunburst。生图使用 POST /v1/images/generations,编辑图片使用 POST /v1/images/edits。
2. Skill 需要同时支持“根据提示词生成图片”和“上传已有图片后按提示词编辑图片”,并把返回的 data[0].b64_json 正确解码为图片文件。
3. 默认 API 地址为 https://waihub.top。直接从 ~/.codex/auth.json 读取 OPENAI_API_KEY 作为 WAIHUB_API_KEY,不需要检查 WAIHUB_API_KEY 是否缺失。
4. 不要在 Skill、脚本、日志或回复中写入、打印或回显真实密钥。
5. 提供清晰的 SKILL.md、可复用的调用脚本和必要的参考说明;自动判断 macOS/Linux/Windows 的 Base64 解码方式,并在写入前检查接口错误。
6. 支持并暴露以下可选参数:`model`、`prompt`、`size`、`quality`、`background`、`output_format`、`output_compression`、`n`、`moderation`、`stream`、`partial_images`、`user`;编辑时还支持 `image`、`mask`、`input_fidelity`。未指定时使用 `model: gpt-image-2.5`、`size: 1024x1024`、`quality: xhigh`、`output_format: png`、`n: 1`。
7. 根据任务自动选择参数:壁纸优先横向尺寸(如 `1536x1024` 或合法自定义尺寸),透明素材使用 `background: transparent` 与 PNG/WebP;Flare 用于快速探索和草稿,Sunburst 用于精修、局部编辑和保持细节。调用前校验枚举值、尺寸限制、压缩参数和编辑输入,发现不兼容时返回清晰错误。
8. 创建完成后校验 Skill 目录结构和脚本语法,并告诉我如何用自然语言触发它。除非我明确要求,不要发起会消耗额度的真实生图请求。

如果 AI 无法访问链接,可以把本页后面的接口示例一并复制给它。创建完成后,就可以直接说“生成一张产品海报”或“把这张图的背景改成蓝色”来调用 Skill。

Waihub 提供 OpenAI 兼容的图像接口。当前 gpt-image-2.5 是默认模型;需要更高级效果时,可以选择 gpt-image-2.5-flare 或 gpt-image-2.5-sunburst。三者都推荐直接使用 /v1/images/* 路由:生成和编辑的返回结构一致,图片内容位于 data[0].b64_json。

Flare 和 Sunburst 怎么选? ​

你的任务建议先选模型重点检查
尝试多个视觉方向Flare构图是否符合需求
制作日常社交配图或产品图草稿Flare文字、物体形状,以及不同版本的一致性
精修已选定的宣传图Sunburst局部修改是否保留已经确认的部分
编辑细节丰富的产品照片Sunburst标签形状、纹理、边缘,以及是否出现意外改动

这个选择建议基于 OpenAI 对 Flare 和 Sunburst 的官方定位。两者都接受文字和图片输入,支持 low、medium、high、xhigh、max 和 auto 质量设置。建议用最难处理的图片比较高质量设置的效果,再决定是否值得使用,无需每张草稿都开到最高。

图像接口参数 ​

以下参数来自 OpenAI Image generation 官方指南 和 Images API Reference。Waihub 的 /v1/images/generations 与 /v1/images/edits 采用兼容字段;如果某个令牌或上游模型尚未开放某个字段,应以接口返回的错误为准。

参数可选值或格式说明
modelgpt-image-2.5、gpt-image-2.5-flare、gpt-image-2.5-sunburst选择图像模型。文档默认使用 gpt-image-2.5;Flare 适合快速探索,Sunburst 适合精修和编辑。
prompt文本字符串描述要生成的画面,或描述对输入图片要执行的编辑。
size1024x1024、1536x1024、1024x1536、auto 或 WIDTHxHEIGHT自定义尺寸时宽高必须是 16 的倍数,宽高比为 1:3 至 3:1,单边不超过 3840 像素,总像素数为 655,360 至 8,294,400;超过 2560x1440 属于实验性范围。
qualitylow、medium、high、xhigh、max、auto控制渲染质量、延迟和成本。Flare/Sunburst 支持 xhigh、max;本文默认 xhigh,官方默认值为 auto。
backgroundauto、opaque、transparent设置背景策略。透明背景需同时使用 PNG 或 WebP。
output_formatpng、jpeg、webp设置输出格式;PNG 是默认格式,JPEG 通常更快。
output_compression0–100JPEG/WebP 的压缩等级;PNG 不使用此参数。
n正整数一次请求生成的图片数量;省略时返回一张。
moderationauto、low调整内容审核严格程度,默认 auto;仍会执行内容政策过滤。
streamtrue、false是否流式接收生成过程。
partial_images0–3流式生成时请求的中间图片数量;设为 0 只返回最终图。
user字符串标识为请求附加调用方用户标识;不要放入密码或令牌。
image一个或多个图片文件/v1/images/edits 的输入图片,也可作为参考图。
mask与输入图同尺寸且含 alpha 通道指定可编辑区域;图片和 mask 的格式、尺寸必须一致,mask 小于 50MB。
input_fidelity由上游模型支持时使用编辑时控制输入图片细节保留程度;被拒绝时删除该字段后重试。

Responses API 使用同一组图像输出选项,但字段放在 tools[] 的图像生成工具配置中;顶层 model 必须是可调用图像工具的编排模型,图像模型写在工具的 model 字段。工具还支持 action: "auto" | "generate" | "edit"。

gpt-image-2.5 效果图

效果图使用的 Prompt ​

本图使用 gpt-image-2.5-flare、quality: xhigh 和横向尺寸生成,实际提示词如下:

text
Create a beautifully composed 16:9 widescreen computer wallpaper of the principal characters from the anime Tokyo Ghoul. Recognizable canonical anime character designs: white-haired Ken Kaneki at the center foreground wearing a black high-collar outfit with a small iconic black ghoul mask, Touka Kirishima with her asymmetric deep indigo bob beside him, silver-haired Rize Kamishiro with glasses behind them, and dark-haired Hideyoshi Nagachika with orange-blond hair plus white-haired Juuzou Suzuya as a small supporting pair. Exactly five characters; Hide is orange-blond, not black-haired. Arrange the ensemble as one calm sculptural group in the lower right two-thirds, varied shoulder heights, composed quiet expressions, elegant three-quarter poses. Keep the upper and left third spacious for desktop icons. Sophisticated minimal anime editorial illustration, large simple color masses, low visual frequency, restrained edge density, clear edge hierarchy, minimal internal contour lines, broad shadow shapes, suppress micro-texture and micro-contrast, no unnecessary specular highlights. Matte warm ivory background occupying most of the image, charcoal-black clothing masses, muted slate-indigo accents, a single restrained dark crimson circular sun behind the group and one broad abstract crimson kagune ribbon used as a clean compositional shape. Use only a tightly curated palette of ivory, near-black charcoal, slate indigo, muted crimson, pale skin. Flat cel shading with only one broad shadow shape per surface, simplified hair grouped into large locks without individual strands, precise sparse facial features, clean distinctive outer silhouettes, no tiny costume details. Strong negative space, tasteful asymmetry, tranquil moody elegance, exceptionally clean wallpaper design. No text, no typography, no logos, no watermark, no border, no city detail, no particles, no splatters, no grain, no hatching, no texture, no glows, no busy background. Fill the entire wide canvas.

密钥安全

示例从 ~/.codex/auth.json 读取 OPENAI_API_KEY 并赋给 WAIHUB_API_KEY。不要把真实密钥提交到代码仓库、文档或日志;如果密钥已经公开,请在令牌管理中轮换。

需要更高级效果时,只需将请求中的 model 替换为 gpt-image-2.5-flare 或 gpt-image-2.5-sunburst,并按需要调整 quality。

例如,横向电脑壁纸可以使用 size: "1536x1024";透明素材可以组合 background: "transparent"、output_format: "png";需要流式预览时可以设置 stream: true 和 partial_images: 2。JPEG/WebP 输出还可以增加 output_compression(0–100)。

成功响应的关键字段如下:

json
{
  "data": [
    { "b64_json": "iVBORw0KGgo..." }
  ]
}

如果响应中有 error,先查看错误内容,不要继续解码空字符串:

bash
jq -c '.error // empty' image_response.json

/v1/edits 是兼容别名,推荐使用完整的 /v1/images/edits。如果上游模型支持多图输入,可以重复传递 image[] 字段。

继续使用 Responses 格式 ​

/v1/responses 仍可用于需要 Responses 外层协议的场景。顶层 model 是负责编排的文本模型,图像模型放在 tools[].model:

bash
curl --fail-with-body --location 'https://waihub.top/v1/responses' \
  --header "Authorization: Bearer ${WAIHUB_API_KEY}" \
  --header 'Content-Type: application/json' \
  --data '{
    "model": "gpt-5.6-terra",
    "input": "画一只可爱的猫抱着水獭",
    "tools": [
      {
        "type": "image_generation",
        "model": "gpt-image-2.5",
        "output_format": "png",
        "quality": "xhigh"
      }
    ],
    "tool_choice": "auto",
    "stream": false,
    "store": false
  }' \
  -o response.json

jq -e -r '
  .output[]?
  | select(.type == "image_generation_call" and .status == "completed")
  | .result
' response.json | base64 -D > cat_otter.png

Responses 的图片字段是 output[].result,不要和 /v1/images/generations 的 data[].b64_json 混用。

常见问题 ​

  • /v1/images/generations 和 /v1/images/edits 使用 data[0].b64_json 提取图片。
  • /v1/responses 使用 output[].result 提取图片。
  • Responses 的顶层模型不要改成图像模型;需要指定图像模型时设置 tools[0].model,默认使用 gpt-image-2.5。
  • macOS 的 base64 解码参数是 -D,Linux 通常是 -d。
  • 可先调用 GET https://waihub.top/v1/models 确认当前令牌能看到 gpt-image-2.5、gpt-image-2.5-flare 或 gpt-image-2.5-sunburst。

Waihub Documentation