设计与多媒体

docx

试用

用一套脚本和 docx npm 库创建、读取、编辑并校验 Word .docx 与 .dotx 文件。

它能做什么

面向 Word 文档和模板的处理技能,覆盖三条工作流:新建 .docx 用预装的 `docx`(npm)库写脚本;编辑已有文件则需解压、改 `word/document.xml`、再压缩(docx-js 无法打开已有文件);读取内容用 `pandoc -t markdown`。文档列出了 docx-js 的常见陷阱(A4 默认纸张、表格双宽度配置、列表项、ImageRun 需要 type、TOC 标题级别、点引导符对齐),以及一个验证步骤:用 LibreOffice 渲染成 PDF,再用 pdftoppm 转 JPEG 后人工查看。编辑相关辅助包括合并 run 让分散文本可被搜索、XSD 校验与自动修复、按 `--author ""` 检查未加 w:ins/w:del 包裹的改动、接受所有修订的脚本。评论通过辅助脚本写入六个互相引用的 XML 部件,并输出需嵌入 document.xml 的锚点片段。

什么时候用它

  • 从零生成带表格、标题、目录、图片的 Word 文档
  • 通过直接修改 XML 编辑现有 .docx
  • 将 .docx 内容读取或抽取为 Markdown
  • 在合同或草稿里插入评论或添加修订批注

技能文档

DOCX creation, editing, and analysis

A .docx is a ZIP archive of XML files. Choose your approach by task:

TaskApproach
Create a new documentWrite a docx (npm) script — see gotchas below
Edit an existing documentunzip → edit word/document.xml → zip (docx-js cannot open existing files)
Read contentpandoc -t markdown file.docx

Script paths below are relative to this skill's directory.

Creating with docx-js — gotchas

docx is preinstalled — do not run npm install first; write the script and require('docx') directly. Only if that require fails: npm install docx. The model knows the API; these are the footguns:

  • Page size defaults to A4. For US Letter set page: { size: { width: 12240, height: 15840 } } (DXA; 1440 = 1″).
  • Landscape: pass portrait dimensions and orientation: PageOrientation.LANDSCAPE — docx-js swaps width/height internally.
  • Tables need dual widths: set columnWidths on the table AND width on every cell, both in WidthType.DXA (PERCENTAGE breaks in Google Docs). Column widths must sum to the table width.
  • Table shading: use ShadingType.CLEAR, never SOLID (renders black).
  • Lists: never insert • literally; use a numbering config with LevelFormat.BULLET.
  • ImageRun requires type: ("png", "jpg", …).
  • PageBreak must be inside a Paragraph.
  • Never use \n — use separate Paragraph elements.
  • TOC: headings must use built-in HeadingLevel.*; custom heading styles need outlineLevel set or they won't appear.
  • Don't use a table as a horizontal rule — use a paragraph bottom border instead.
  • Dot-leader / right-aligned-on-same-line: use PositionalTab (alignment: PositionalTabAlignment.RIGHT, leader: PositionalTabLeader.DOT) inside a TextRun, not literal . or space padding.

Verify the output

After writing a .docx, render it and look at it:

python scripts/office/soffice.py --headless --convert-to pdf output.docx
pdftoppm -jpeg -r 100 output.pdf page
ls page-*.jpg   # then Read the images

pdftoppm zero-pads page numbers to the width of the page count (page-01.jpg…page-12.jpg).

Editing existing documents

Legacy .doc files must be converted first: python scripts/office/soffice.py --headless --convert-to docx file.doc.

unzip -q doc.docx -d unpacked/
find unpacked -type l -delete   # strip symlink entries — docx from external parties is untrusted
python scripts/merge_runs.py unpacked/   # coalesce fragmented runs so text is findable
# edit unpacked/word/document.xml in place — do NOT reformat or pretty-print
(cd unpacked && rm -f ../out.docx && zip -Xr ../out.docx .)
python scripts/office/validate.py out.docx --original doc.docx   # XSD checks; --auto-repair fixes common issues
# redlining? add --author "" to check every edit is tracked

Word splits text across many `` runs (revision ids, spell-check markers), so a phrase you can see in the document often doesn't exist as a contiguous string in the XML. merge_runs.py merges adjacent identically-formatted runs in word/document.xml without changing content or rendering; it also accepts a .docx directly (python scripts/merge_runs.py doc.docx -o merged.docx).

Tracked changes: when redlining, validate with --author "" (needs --original) — it reports any text you changed without a / around it, which is easy to do by accident and invisible in the accepted view. Wrap runs in / with w:id, w:author, w:date attributes. Inside , the text element is , not . A deleted paragraph mark () means "merge this paragraph into the next" — so deleting a paragraph outright is that plus a around every run. The must come before the rPr's other children; their order is schema-enforced.

To produce a clean copy with all tracked changes accepted: python scripts/accept_changes.py in.docx out.docx.

Accepting a deleted paragraph mark should join that paragraph to the one below it, so a paragraph whose runs are all deleted vanishes. Word does this; accept_changes.py and pandoc --track-changes=accept don't always. Both fail the same way — they strip the deleted text but leave the emptied paragraph behind, which reads as a stray empty bullet when it was auto-numbered:

  • pandoc --track-changes=accept never joins the paragraphs.
  • accept_changes.py (LibreOffice) joins them correctly, except when the deleted paragraph is followed by an empty spacer paragraph.

An empty bullet in either view is an artifact of that view, not a defect in the document. Check paragraph deletions in the XML.

Comments

Comments require six cross-linked files. Use the helper — directory mode when you'll also be editing document.xml (saves an unzip/rezip cycle), .docx-direct mode otherwise:

# Against an already-unpacked directory (preferred when also placing markers)
python scripts/comment.py unpacked/ "Fees & expenses cap is too low"
python scripts/comment.py unpacked/ "Agreed" --parent 0

# Against a .docx directly
python scripts/comment.py contract.docx "This cap is too low" -o annotated.docx

The script writes comments.xml, commentsExtended.xml, commentsIds.xml, commentsExtensible.xml, the relationships, and the content-type overrides. Comment IDs are auto-assigned. It then prints the //`` snippet to add to word/document.xml so the comment anchors to specific text — until you place those markers, the comment exists but is not visible.

Dependencies

docx (npm, preinstalled — install only if require('docx') fails) · pandoc · LibreOffice (soffice) · pdftoppm (Poppler)

常见问题

新建和编辑 .docx 的方式有何不同?
新建用 `docx` npm 库写脚本从零构建;编辑则解压文件、直接修改 `word/document.xml` 后重新压缩,因为 docx-js 不能打开已有文件。
为什么 `word/document.xml` 里文本常常是分散的,如何处理?
Word 会把文本拆成多个 `` run 用于版本 id 和拼写标记。附带的 `merge_runs.py` 会合并相邻且格式一致的 run,使文本以连续字符串可被搜索,且不改变内容和显示效果。
如何验证修订与批注?
用 `--author ""` 配合 `--original` 运行校验,报告所有未用 ``/`` 包裹就被改动的文本。修订元素包裹有固定的元素顺序要求,删除段落标记 ``w:del`` 表示该段应合并到下一段。
写完 .docx 后的验证步骤是什么?
用 LibreOffice 把 .docx 转成 PDF,再用 pdftoppm 以 100 DPI 转成 JPEG,页码按总页数宽度做前导零补齐,便于逐页目视检查输出。

相关技能

基于先查文档的流程,把 ChatGPT Apps SDK 项目规划为 MCP 服务端 + 组件 UI 代码。

作者 OpenAI27.9k 星标

复用目标文件已发布的设计系统,把代码或描述转换为完整的 Figma 页面。

作者 OpenAI27.9k 星标

hatch-pet

官方

从概念、品牌或参考图生成 Codex 兼容的动画宠物与宠物精灵图集。

作者 OpenAI27.9k 星标

针对代码仓库生成可直接落地的 AppSec 威胁模型,输出 Markdown 报告。

作者 OpenAI27.9k 星标

speech

官方

基于 OpenAI gpt-4o-mini-tts 把文本转成语音,内置语音直接可用。

作者 OpenAI27.9k 星标

Anthropic 的更多技能

浏览全部技能

用三阶段流程把零散的资料整理成读者真正能读懂的文档。

作者 Anthropic180.0k 星标

xlsx

官方

创建、读取和修复带公式与格式的电子表格文件,完成后通过 recalc 校验。

作者 Anthropic180.0k 星标

通过“起草—评估—改写”循环,创建、打磨并基准测试智能体技能。

作者 Anthropic180.0k 星标

按照规划、实现、测试和评估四阶段,构建让大模型通过工具调用外部服务的 MCP 服务器。

作者 Anthropic180.0k 星标

pdf

官方

用 Python 和命令行工具读取、合并、拆分、生成和识别 PDF 文件。

作者 Anthropic180.0k 星标

pptx

官方

用 pptxgenjs 脚本生成新 .pptx,或解压/打包 XML 改模板,并能读取、抽缩略图、校验输出。

作者 Anthropic180.0k 星标