docx
OfficialCreate, read, and edit Word .docx and .dotx files from scripts.
Design & media
Transcribe audio files to text with optional diarization and known-speaker hints. Use when a user asks to transcribe speech from audio/video, extract text from recordings, or label speakers in intervi
Transcribe audio files to text with optional diarization and known-speaker hints. Use when a user asks to transcribe speech from audio/video, extract text from recordings, or label speakers in interviews or meetings.
Transcribe audio using OpenAI, with optional speaker diarization when requested. Prefer the bundled CLI for deterministic, repeatable runs.
OPENAI_API_KEY is set. If missing, ask the user to set it locally (do not ask them to paste the key).transcribe_diarize.py CLI with sensible defaults (fast text transcription).output/transcribe/ when working in this repo.gpt-4o-mini-transcribe with --response-format text for fast transcription.--model gpt-4o-transcribe-diarize --response-format diarized_json.--chunking-strategy auto.gpt-4o-transcribe-diarize.output/transcribe// for evaluation runs.--out-dir for multiple files to avoid overwriting.Prefer uv for dependency management.
uv pip install openai
If uv is unavailable:
python3 -m pip install openai
OPENAI_API_KEY must be set for live API calls.export CODEX_HOME="${CODEX_HOME:-$HOME/.codex}"
export TRANSCRIBE_CLI="$CODEX_HOME/skills/transcribe/scripts/transcribe_diarize.py"
User-scoped skills install under $CODEX_HOME/skills (default: ~/.codex/skills).
Single file (fast text default):
python3 "$TRANSCRIBE_CLI" \
path/to/audio.wav \
--out transcript.txt
Diarization with known speakers (up to 4):
python3 "$TRANSCRIBE_CLI" \
meeting.m4a \
--model gpt-4o-transcribe-diarize \
--known-speaker "Alice=refs/alice.wav" \
--known-speaker "Bob=refs/bob.wav" \
--response-format diarized_json \
--out-dir output/transcribe/meeting
Plain text output (explicit):
python3 "$TRANSCRIBE_CLI" \
interview.mp3 \
--response-format text \
--out interview.txt
references/api.md: supported formats, limits, response formats, and known-speaker notes.Create, read, and edit Word .docx and .dotx files from scripts.
Scaffold ChatGPT Apps SDK projects with a docs-first workflow, tool plan, and working MCP server plus widget code.
Build a complete, code-aligned Figma design system in phases, with variables, components, and themes created in the right order.
Build Codex-compatible animated pets and 8x9 sprite atlases from text, references, or brand cues.
Bootstrap, scaffold, and verify WinUI 3 desktop apps with C# and the Windows App SDK.
Connect Figma design components to their code counterparts via Code Connect mappings.
Scaffold ChatGPT Apps SDK projects with a docs-first workflow, tool plan, and working MCP server plus widget code.
Build a complete, code-aligned Figma design system in phases, with variables, components, and themes created in the right order.
Write safe, incremental JavaScript against the Figma Plugin API from your agent.
Build Codex-compatible animated pets and 8x9 sprite atlases from text, references, or brand cues.
Generate or edit raster images via a built-in tool, with a CLI fallback only on explicit request.
Get authoritative, current answers from OpenAI developer docs with citations and source paths.