用三阶段流程把零散的资料整理成读者真正能读懂的文档。
文档
xlsx
试用创建、读取和修复带公式与格式的电子表格文件,完成后通过 recalc 校验。
它能做什么
基于 openpyxl 与 pandas,对 .xlsx/.xlsm/.xltx/.csv/.tsv 文件进行读取、编辑、创建与格式转换。输出遵循统一规范(Arial/Times New Roman 字体、公式驱动单元格、注释说明每个假设),并通过强制的 `recalc.py` 步骤在交付前捕获 #NAME?、#REF!、#DIV/0! 等错误。同时包含财务模型约定(输入与公式的颜色编码、货币/百分比格式、零显示为短横线)以及在 LibreOffice 重算引擎下的函数兼容规则。
什么时候用它
- 打开并编辑已有的 .xlsx 或 .xlsm 文件
- 构建带公式与格式的财务模型
- 用 pandas 批量读写表格数据
- 在 .xlsx、.csv、.tsv 之间转换格式
技能文档
XLSX creation, editing, and analysis
| Task | Approach |
|---|---|
| Create or edit with formulas/formatting | openpyxl — see gotchas below |
| Bulk data in or out | pandas (read_excel, to_excel) |
| Quick look at a sheet | markitdown file.xlsx — ## SheetName per sheet; reads .xlsm too. No cell coordinates, so don't plan edits from it |
| Read a model (formulas and values) | two load_workbook passes — see gotchas |
openpyxl,pandas, andmarkitdownare preinstalled — do not runpip installfirst; write the script and import directly. Only if an import fails (or themarkitdowncommand is missing):pip installthe missing package.
Script paths below are relative to this skill's directory.
Requirements for every output
- Professional font (Arial, Times New Roman) throughout, unless the user says otherwise.
- Zero formula errors. Never ship while
recalc.pyreportserrors_found. If you think an error predates you, prove it: load the original withdata_only=Trueand look at that cell. An error you introduced looks exactly like one you inherited. - Use formulas, never hardcoded results. Write
sheet['B10'] = '=SUM(B2:B9)', not the Python-computed total. The sheet must recalculate when its inputs change. - Follow the user's spec literally. Exact tab names, exact column headers, and the formula they spelled out. A redesign that computes something else fails, however elegant.
- Document every assumption and hardcoded number where the reader will see it — a cell comment, or an adjacent cell at a table's end. Cite a real source when one exists (
Source: Company 10-K, FY2024, Page 45, Revenue Note, [SEC EDGAR URL]); when the number came from the user, say so plainly. - A workbook you create for someone to fill in needs a short legend naming which cells to edit, and one example row of realistic values showing the expected format. Never add such a row to a file you were asked to edit.
- Editing an existing file: match its conventions exactly. They override every guideline here. Find its designated input cells first — a distinct font color, fill, or shading marks them — write only there, and leave every existing formula untouched.
Recalculate (mandatory whenever the file contains formulas)
openpyxl writes formulas as strings with no cached values. Until you recalculate, every
formula cell reads back as None to anything reading cached values — pandas,
load_workbook(data_only=True), and most previewers.
python scripts/recalc.py output.xlsx [timeout_seconds] # default 30
LibreOffice computes every formula, the file is rewritten in place, and you get JSON:
status (success | errors_found), total_formulas, total_errors, and an
error_summary naming up to 100 cells per error type (locations_truncated says how many it
withheld — trust total_errors, not the length of the list). Fix what it names and run it
again. JSON with an error key instead of a status means nothing was recalculated, and
only that case exits non-zero — errors_found exits 0, so never treat a clean exit as a clean
workbook.
A green recalc proves your formulas evaluate, not that they are right. An off-by-one range or a reference to the wrong row yields a clean, error-free file with wrong numbers. Write 2–3 formulas first and check they pull the values you expect, before building out a grid.
A workbook that links to another file loses those links if you re-save it with openpyxl and
then recalculate. Such a formula reads ='[1]Returns Analysis'!$B$2 — the [1] is an index
into the workbook's external-reference list, naming a separate file on disk, not a sheet.
That file is rarely present here, so the cell's cached value is the only thing holding its
data. openpyxl strips that value on save; LibreOffice then has to resolve the reference for
real, fails, writes #NAME?, and deletes every link. recalc.py refuses to run in that state
— copy those cells' values out of the original before you save over them (--force overrides,
and accepts the loss).
Choosing formulas that survive verification
LibreOffice implements fewer functions than Excel, and one it cannot evaluate becomes a
literal #NAME? baked into the file you deliver.
- Prefer Excel-2007-era functions —
SUMIFS,INDEX,MATCH,IFERROR,SUMPRODUCT— which need no prefix. - Six post-2007 functions work, but only with an
_xlfn.prefix, because openpyxl writes your formula into the XML verbatim and Excel stores post-2007 names prefixed (its UI hides the prefix):_xlfn.TEXTJOIN,_xlfn.CONCAT,_xlfn.IFS,_xlfn.SWITCH,_xlfn.MAXIFS,_xlfn.MINIFS. Written bare, each yields#NAME?. - Never use
XLOOKUP,XMATCH,SORT,FILTER,UNIQUE, orSEQUENCE. The runtime's LibreOffice cannot evaluate them under any prefix. Newer builds do evaluate them, but they are spilling array functions and an openpyxl-written file has no spill metadata, so only the top-left cell of the range gets a value — andrecalc.pyreportstotal_errors: 0on the truncated result. UseINDEX/MATCHfor lookups, and sort, filter, and de-duplicate in Python before writing the cells. - A formula LibreOffice could not parse is written back lowercased — a quick tell beside a
#NAME?.
openpyxl gotchas
- Reading a model takes two loads.
data_only=Trueyields cached values with the formulas gone; the default yields formula strings with no values. One pass cannot give you both. data_only=Trueis destructive if you save. That workbook has no formulas left, so saving replaces every one with a literal — permanently.data_only=Trueon a file openpyxl just wrote returnsNoneeverywhere — runrecalc.pyfirst. (A formula whose result is""also reads back asNone.)- Merged cells: write the top-left anchor only. Every other cell in the range is a
MergedCellwhose.valueis read-only. .xlsmloses its macros unless you passkeep_vba=Truetoload_workbook.- A sheet name containing a space must be quoted in a cross-sheet reference:
='Assumptions Inputs'!$B$5. Unquoted, it evaluates to#VALUE!.
Financial models
Unless the user says otherwise, or the existing file already does something else.
Color: blue text (0,0,255) for hardcoded inputs and scenario levers · black for formulas ·
green (0,128,0) for links to another sheet · red (255,0,0) for links to another file ·
yellow fill (255,255,0) for key assumptions and cells the user should fill in.
Numbers: currency $#,##0, with the unit named in the header (Revenue ($mm)) · zeros
render as -, including in percentages ($#,##0;($#,##0);-) · negatives in parentheses ·
percentages 0.0%, stored as fractions (0.15 renders 15.0%; storing 15 renders
1500.0%) · valuation multiples 0.0x · years as text ("2024", never 2,024).
Structure: every assumption in its own labeled cell, referenced by the formulas that use it
(=B5*(1+$B$6), never =B5*1.05) · formulas consistent across every projection period, since a
lone edited cell mid-row is the commonest silent error · guard denominators that can be zero.
Dependencies
openpyxl, pandas, markitdown (pip, preinstalled — install only if an import fails or the command is missing) · LibreOffice (soffice, auto-configured for sandboxed environments via scripts/office/soffice.py)
常见问题
- 编辑后为什么必须运行 recalc.py?
- openpyxl 仅把公式以字符串形式写入文件,缓存值为空。因此依赖缓存值的读取方式(pandas、带 data_only=True 的 load_workbook、多数预览工具)都会得到 None,直到 LibreOffice 在原文件上重算为止。
- 我的公式为什么变成 #NAME??
- LibreOffice 支持的函数少于 Excel。2007 年之后新增的函数需要加 _xlfn. 前缀(TEXTJOIN、CONCAT、IFS、SWITCH、MAXIFS、MINIFS);而 XLOOKUP、XMATCH、SORT、FILTER、UNIQUE、SEQUENCE 在任何前缀下都不被支持,且不会触发 total_errors,只会静默截断。
- 能否一次同时读取公式和值?
- 不能。load_workbook(data_only=True) 只返回缓存值、不带公式;默认加载只返回公式字符串、不带值。必须分两次加载,并且需要先运行 recalc.py,否则 data_only=True 全部返回 None。
相关技能
基于 OpenAI 官方开发者文档,给出带引用的权威解答。
获得针对具体语言和框架的安全审查,输出按严重程度排好级的报告与修复建议。
pptx
官方用 pptxgenjs 脚本生成新 .pptx,或解压/打包 XML 改模板,并能读取、抽缩略图、校验输出。
Use when a user asks to debug or fix failing GitHub PR checks that run in GitHub Actions; use `gh` to inspect checks and logs, summarize failure context, draft a fix plan, and implement only after exp
Toolkit for styling artifacts with a theme. These artifacts can be slides, docs, reportings, HTML landing pages, etc. There are 10 pre-set themes with colors/fonts that you can apply to any artifact t
Anthropic 的更多技能
浏览全部技能用三阶段流程把零散的资料整理成读者真正能读懂的文档。
通过“起草—评估—改写”循环,创建、打磨并基准测试智能体技能。
按照规划、实现、测试和评估四阶段,构建让大模型通过工具调用外部服务的 MCP 服务器。
docx
官方用一套脚本和 docx npm 库创建、读取、编辑并校验 Word .docx 与 .dotx 文件。
用 Python 和命令行工具读取、合并、拆分、生成和识别 PDF 文件。
pptx
官方用 pptxgenjs 脚本生成新 .pptx,或解压/打包 XML 改模板,并能读取、抽缩略图、校验输出。