基于先查文档的流程,把 ChatGPT Apps SDK 项目规划为 MCP 服务端 + 组件 UI 代码。
设计与多媒体
transcribe
试用Transcribe audio files to text with optional diarization and known-speaker hints. Use when a user asks to transcribe speech from audio/video, extract text from recordings, or label speakers in intervi
它能做什么
Transcribe audio files to text with optional diarization and known-speaker hints. Use when a user asks to transcribe speech from audio/video, extract text from recordings, or label speakers in interviews or meetings.
技能文档
Audio Transcribe
Transcribe audio using OpenAI, with optional speaker diarization when requested. Prefer the bundled CLI for deterministic, repeatable runs.
Workflow
- Collect inputs: audio file path(s), desired response format (text/json/diarized_json), optional language hint, and any known speaker references.
- Verify
OPENAI_API_KEYis set. If missing, ask the user to set it locally (do not ask them to paste the key). - Run the bundled
transcribe_diarize.pyCLI with sensible defaults (fast text transcription). - Validate the output: transcription quality, speaker labels, and segment boundaries; iterate with a single targeted change if needed.
- Save outputs under
output/transcribe/when working in this repo.
Decision rules
- Default to
gpt-4o-mini-transcribewith--response-format textfor fast transcription. - If the user wants speaker labels or diarization, use
--model gpt-4o-transcribe-diarize --response-format diarized_json. - If audio is longer than ~30 seconds, keep
--chunking-strategy auto. - Prompting is not supported for
gpt-4o-transcribe-diarize.
Output conventions
- Use
output/transcribe//for evaluation runs. - Use
--out-dirfor multiple files to avoid overwriting.
Dependencies (install if missing)
Prefer uv for dependency management.
uv pip install openai
If uv is unavailable:
python3 -m pip install openai
Environment
OPENAI_API_KEYmust be set for live API calls.- If the key is missing, instruct the user to create one in the OpenAI platform UI and export it in their shell.
- Never ask the user to paste the full key in chat.
Skill path (set once)
export CODEX_HOME="${CODEX_HOME:-$HOME/.codex}"
export TRANSCRIBE_CLI="$CODEX_HOME/skills/transcribe/scripts/transcribe_diarize.py"
User-scoped skills install under $CODEX_HOME/skills (default: ~/.codex/skills).
CLI quick start
Single file (fast text default):
python3 "$TRANSCRIBE_CLI" \
path/to/audio.wav \
--out transcript.txt
Diarization with known speakers (up to 4):
python3 "$TRANSCRIBE_CLI" \
meeting.m4a \
--model gpt-4o-transcribe-diarize \
--known-speaker "Alice=refs/alice.wav" \
--known-speaker "Bob=refs/bob.wav" \
--response-format diarized_json \
--out-dir output/transcribe/meeting
Plain text output (explicit):
python3 "$TRANSCRIBE_CLI" \
interview.mp3 \
--response-format text \
--out interview.txt
Reference map
references/api.md: supported formats, limits, response formats, and known-speaker notes.
相关技能
复用目标文件已发布的设计系统,把代码或描述转换为完整的 Figma 页面。
从概念、品牌或参考图生成 Codex 兼容的动画宠物与宠物精灵图集。
针对代码仓库生成可直接落地的 AppSec 威胁模型,输出 Markdown 报告。
为 AI 编码代理生成针对你项目定制的 Figma 设计系统规则文件。
speech
官方基于 OpenAI gpt-4o-mini-tts 把文本转成语音,内置语音直接可用。
OpenAI 的更多技能
浏览全部技能基于先查文档的流程,把 ChatGPT Apps SDK 项目规划为 MCP 服务端 + 组件 UI 代码。
复用目标文件已发布的设计系统,把代码或描述转换为完整的 Figma 页面。
通过 `use_figma` MCP 工具在 Figma 文件中执行 JavaScript 的参考技能,定义插件代码的 API 规则与写法。
从概念、品牌或参考图生成 Codex 兼容的动画宠物与宠物精灵图集。
imagegen
官方在 Codex 项目里直接生成或编辑位图图像,并把项目用到的素材放回项目内
基于 OpenAI 官方开发者文档,给出带引用的权威解答。