设计与多媒体

docx

试用

通过 docx-js、原始 XML 和辅助脚本创建、读取与编辑 .docx 与 .dotx 文件。

它能做什么

.docx 本质上是 XML 的 ZIP 压缩包。该技能覆盖三类工作流:用预装的 docx(npm)脚本生成新文件(处理页大小、表格、目录、列表、图片、定位制表符);通过解压、修改 word/document.xml、再压缩来编辑现有文件;用 pandoc 把内容读出来。技能内置的辅助脚本可合并被碎片化的 run、对 XML 做校验与自动修复、在六个相互关联的文件中插入批注、以及接受所有修订变更。旧版 .doc 文件需先用 LibreOffice 转成 .docx。生成后用 soffice 转 PDF,再用 pdftoppm 渲染成 JPEG 检查每一页,再交付。

什么时候用它

  • 生成带标题、表格和图片的 Word 文档
  • 在已有 .docx 中修改文本或结构
  • 把 .docx 内容抽取或转换为 Markdown
  • 添加修订或批注,并产出已接受所有修订的版本

技能文档

DOCX creation, editing, and analysis

A .docx is a ZIP archive of XML files. Choose your approach by task:

TaskApproach
Create a new documentWrite a docx (npm) script — see gotchas below
Edit an existing documentunzip → edit word/document.xml → zip (docx-js cannot open existing files)
Read contentpandoc -t markdown file.docx

Script paths below are relative to this skill's directory.

Creating with docx-js — gotchas

docx is preinstalled — do not run npm install first; write the script and require('docx') directly. Only if that require fails: npm install docx. The model knows the API; these are the footguns:

  • Page size defaults to A4. For US Letter set page: { size: { width: 12240, height: 15840 } } (DXA; 1440 = 1″).
  • Landscape: pass portrait dimensions and orientation: PageOrientation.LANDSCAPE — docx-js swaps width/height internally.
  • Tables need dual widths: set columnWidths on the table AND width on every cell, both in WidthType.DXA (PERCENTAGE breaks in Google Docs). Column widths must sum to the table width.
  • Table shading: use ShadingType.CLEAR, never SOLID (renders black).
  • Lists: never insert • literally; use a numbering config with LevelFormat.BULLET.
  • ImageRun requires type: ("png", "jpg", …).
  • PageBreak must be inside a Paragraph.
  • Never use \n — use separate Paragraph elements.
  • TOC: headings must use built-in HeadingLevel.*; custom heading styles need outlineLevel set or they won't appear.
  • Don't use a table as a horizontal rule — use a paragraph bottom border instead.
  • Dot-leader / right-aligned-on-same-line: use PositionalTab (alignment: PositionalTabAlignment.RIGHT, leader: PositionalTabLeader.DOT) inside a TextRun, not literal . or space padding.

Verify the output

After writing a .docx, render it and look at it:

python scripts/office/soffice.py --headless --convert-to pdf output.docx
pdftoppm -jpeg -r 100 output.pdf page
ls page-*.jpg   # then Read the images

pdftoppm zero-pads page numbers to the width of the page count (page-01.jpg…page-12.jpg).

Editing existing documents

Legacy .doc files must be converted first: python scripts/office/soffice.py --headless --convert-to docx file.doc.

unzip -q doc.docx -d unpacked/
find unpacked -type l -delete   # strip symlink entries — docx from external parties is untrusted
python scripts/merge_runs.py unpacked/   # coalesce fragmented runs so text is findable
# edit unpacked/word/document.xml in place — do NOT reformat or pretty-print
(cd unpacked && rm -f ../out.docx && zip -Xr ../out.docx .)
python scripts/office/validate.py out.docx --original doc.docx   # XSD checks; --auto-repair fixes common issues
# redlining? add --author "" to check every edit is tracked

Word splits text across many `` runs (revision ids, spell-check markers), so a phrase you can see in the document often doesn't exist as a contiguous string in the XML. merge_runs.py merges adjacent identically-formatted runs in word/document.xml without changing content or rendering; it also accepts a .docx directly (python scripts/merge_runs.py doc.docx -o merged.docx).

Tracked changes: when redlining, validate with --author "" (needs --original) — it reports any text you changed without a / around it, which is easy to do by accident and invisible in the accepted view. Wrap runs in / with w:id, w:author, w:date attributes. Inside , the text element is , not . A deleted paragraph mark () means "merge this paragraph into the next" — so deleting a paragraph outright is that plus a around every run. The must come before the rPr's other children; their order is schema-enforced.

To produce a clean copy with all tracked changes accepted: python scripts/accept_changes.py in.docx out.docx.

Accepting a deleted paragraph mark should join that paragraph to the one below it, so a paragraph whose runs are all deleted vanishes. Word does this; accept_changes.py and pandoc --track-changes=accept don't always. Both fail the same way — they strip the deleted text but leave the emptied paragraph behind, which reads as a stray empty bullet when it was auto-numbered:

  • pandoc --track-changes=accept never joins the paragraphs.
  • accept_changes.py (LibreOffice) joins them correctly, except when the deleted paragraph is followed by an empty spacer paragraph.

An empty bullet in either view is an artifact of that view, not a defect in the document. Check paragraph deletions in the XML.

Comments

Comments require six cross-linked files. Use the helper — directory mode when you'll also be editing document.xml (saves an unzip/rezip cycle), .docx-direct mode otherwise:

# Against an already-unpacked directory (preferred when also placing markers)
python scripts/comment.py unpacked/ "Fees & expenses cap is too low"
python scripts/comment.py unpacked/ "Agreed" --parent 0

# Against a .docx directly
python scripts/comment.py contract.docx "This cap is too low" -o annotated.docx

The script writes comments.xml, commentsExtended.xml, commentsIds.xml, commentsExtensible.xml, the relationships, and the content-type overrides. Comment IDs are auto-assigned. It then prints the //`` snippet to add to word/document.xml so the comment anchors to specific text — until you place those markers, the comment exists but is not visible.

Dependencies

docx (npm, preinstalled — install only if require('docx') fails) · pandoc · LibreOffice (soffice) · pdftoppm (Poppler)

常见问题

如何编辑已有的 Word 文件?
把 .docx 解压到目录,删除软链接,运行 merge_runs.py 合并被碎片化的 run,原地修改 word/document.xml(不要重新格式化),再用 zip 打包为 out.docx,最后运行 validate.py 校验。
需要手动安装 docx 这个 npm 包吗?
不需要,已经预装。直接写脚本并 require('docx');只有当 require 报错时才运行 npm install docx。
如何确认生成的 .docx 看起来正确?
用 soffice --headless --convert-to pdf 转成 PDF,再用 pdftoppm -jpeg -r 100 渲染为 JPEG,然后逐页查看图像。

相关技能

用文档优先的流程搭建 ChatGPT Apps SDK 项目,产出工具规划、MCP 服务端与 Widget 脚手架。

作者 OpenAI27.9k 星标

通过复用已发布的设计系统(组件、变量、样式)来构建或更新完整的 Figma 页面。

作者 OpenAI27.9k 星标

按正确顺序在 Figma 中搭建与代码对齐的完整设计系统,覆盖变量、组件与主题。

作者 OpenAI27.9k 星标

hatch-pet

官方

从概念、品牌线索或参考图出发,生成符合 Codex 规范的动画宠物图集与打包产物。

作者 OpenAI27.9k 星标

winui-app

官方

用 C# 和 Windows App SDK 引导、搭建并验证 WinUI 3 桌面应用。

作者 OpenAI27.9k 星标

基于代码仓库证据输出的 AppSec 威胁模型,落地为一份简洁的 Markdown 文件。

作者 OpenAI27.9k 星标

Anthropic 的更多技能

浏览全部技能

pptx

官方

创建、编辑和读取 .pptx 与 .potx 文件,自带可复用脚本与校验工具。

作者 Anthropic180.0k 星标

xlsx

官方

读写、编辑和分析 Excel 及表格文件,支持公式、格式与重算。

作者 Anthropic180.0k 星标

通过三阶段流程,把零散的资料打磨成一份经得起读者检验的文档草稿。

作者 Anthropic180.0k 星标

pdf

官方

用 Python 库和命令行工具读写、合并、拆分、旋转、生成和加密 PDF 文件。

作者 Anthropic180.0k 星标

把一段简短提示变成 p5.js 生成艺术,附带可调参数和种子控件。

作者 Anthropic180.0k 星标

通过「起草—测试—评估」的循环,帮你创建、改进并量化考核各项 agent 技能。

作者 Anthropic180.0k 星标