Merge pull request 'Markdownテーブル形式に対応' (#8) from feature/markdown-table-format into main
CI / test (push) Successful in 11s
Release / release (push) Successful in 1m54s

Reviewed-on: #8
This commit was merged in pull request #8.
This commit is contained in:
2026-09-05 16:59:36 +09:00
11 changed files with 284 additions and 20 deletions
+8 -5
View File
@@ -3,7 +3,7 @@
`dataxl` は、ファイル形式ごとの差を小さくするために、内部表現を2つに分けています。
- structured value: JSON/YAML/TOMLから読める `map[string]any`, `[]any`, scalar
- table: Excel/CSV/TSVに近い `Header []string` と `Rows [][]string`
- table: Excel/CSV/TSV/Markdownに近い `Header []string` と `Rows [][]string`
変換は原則として次のどれかです。
@@ -19,7 +19,7 @@
- `main.go`: CLI option、stdin/stdout、ファイル入出力
- `conversion.go`: 変換経路の選択、形式名の正規化・推定
- `structured.go`: JSON/YAML/TOML adapter
- `table.go`: CSV/TSV/XLSX adapter、table model、flatten
- `table.go`: CSV/TSV/XLSX/Markdown adapter、table model、flatten
- `path.go`: セル値の型推定、列パスのparse、unflatten
`convert` はファイル入出力から独立しているため、CLIを経由せず変換matrixをテストできます。
@@ -40,9 +40,12 @@ type table struct {
}
```
CSV/TSV/XLSXの読み込みでは、短い行を空文字で埋めて列数を揃えます。ヘッダーより
CSV/TSV/XLSX/Markdownの読み込みでは、短い行を空文字で埋めて列数を揃えます。ヘッダーより
長い行は、名前のない値を破棄しないようエラーにします。
XLSXの書き出しではヘッダーを太字にし、1行目を固定します。
Markdownでは区切り行のalignment markerを受け付けますが、table modelには保持しません。
セル内のdelimiterとバックスラッシュはbackslash escape、改行と前後空白はHTML文字参照で
可逆化します。
## Flattening
@@ -129,11 +132,11 @@ TOMLはトップレベル配列を直接表せないため、表からTOMLへ出
- `gopkg.in/yaml.v3`: YAML読み書き
- `github.com/BurntSushi/toml`: TOML読み書き
Go 1.25以上を前提にしています。
Go 1.25.13以上を前提にしています。
## Error handling
- 未対応形式、decode失敗、workbook/sheet操作失敗は呼び出し元へerrorを返します。
- CSV/TSVのinvalid UTF-8、headerより長い行、値を持つ空header列を拒否します。
- CSV/TSV/Markdownのinvalid UTF-8、headerより長い行、値を持つ空header列を拒否します。
- path復元中の重複header、構文エラー、型競合は変換全体をerrorにします。
- XLSXのstyle・pane設定も通常の変換errorとして扱い、不完全なworkbookを成功扱いしません。
+2 -1
View File
@@ -2,7 +2,7 @@
## Requirements
- Go 1.25 or later
- Go 1.25.13 or later
## Setup
@@ -58,6 +58,7 @@ The current tests cover:
- YAML -> TSV flattening
- TSV -> JSON path restoration
- JSON -> XLSX -> JSON round trip
- Markdown table parsing, escaping, and all-format conversion round trips
- structured -> structured conversion without CLI/file I/O
- extension normalization and format inference
- short table row padding and wider-row rejection