feat(text chat): expose image input via repeatable --image (multimodal M3, multi-image) - #225
feat(text chat): expose image input via repeatable --image (multimodal M3, multi-image)#225shoemoney wants to merge 1 commit into
Conversation
MiniMax-M3 is multimodal, but `text chat` had no way to send an image —
users had to hand-write a base64 messages JSON file, and the obvious
OpenAI `image_url` shape is rejected because the CLI posts to the
Anthropic-compatible /messages endpoint.
- `--image <path-or-url>` on `text chat`, repeatable, so multi-image
compare/diff works in one call
- new `toImageBlock()` in utils/image reuses `toDataUri()` (local paths,
http(s) URLs, existing data URIs) and emits the Anthropic block shape
`{ type: 'image', source: { type: 'base64', media_type, data } }`
- images append to the last user message, promoting string content to a
block array; with no `--message` they become the user message
- images force `MiniMax-M3` when `--model` is unset, so a text-only
`defaultTextModel` in config can't silently break the request
- docs: README, README_CN, skill/SKILL.md (incl. the OpenAI-vs-Anthropic
block-shape gotcha)
Closes MiniMax-AI#224
NianJiuZst
left a comment
There was a problem hiding this comment.
I reviewed the current head and the feature direction is useful, but I do not recommend merging this revision yet.
[P2] Enforce the M3 image contract before encoding. The current MiniMax Anthropic API docs cap each image at 10 MB and the whole request body at 64 MB, and support JPEG/PNG/GIF/WEBP. The new path delegates to toDataUri(), which leaves local/data-URI inputs unbounded, permits remote images up to 50 MB, rejects local GIF, and accepts HEIC/HEIF. A local 11 MiB .png was encoded to 15,379,116 base64 characters without error; repeated --image inputs can therefore exceed the request cap and fail only after substantial memory/network work. Please add chat-specific per-image/format validation and an aggregate request-size check.
Official contract: https://platform.minimax.io/docs/api-reference/text-anthropic-api
Local verification: typecheck, lint (one pre-existing warning), build, 17 focused tests, and the full suite (455/455) passed. Those mocks do not exercise the oversized/unsupported real-provider cases.
Closes #224.
What
mmx text chatnow takes a repeatable--image <path-or-url>:Why
M3 is multimodal and already accepts multiple images through
--messages-file, but there was no CLI path to it — you had to hand-write a base64 messages JSON. Worse, the documented OpenAI shape ({"type":"image_url", ...}) is rejected, becausetext chatposts to the Anthropic-compatible/anthropic/v1/messagesendpoint and needs{"type":"image","source":{"type":"base64",...}}.--imageemits the right shape for you.How
toImageBlock()insrc/utils/image.tswraps the existingtoDataUri()(local paths,http(s)URLs, and pre-made data URIs all work) and returns an Anthropic image block.src/commands/text/chat.tsappends the blocks to the last user message, promotingcontentfrom a string to a block array. Text goes first so the model reads the instruction before the pixels. With no--message, the images become the user message.ContentBlockgains theimagevariant.--imageis present and--modelis not, the model resolves toMiniMax-M3— otherwise a text-onlydefaultTextModelin config would silently break every image request. An explicit--modelstill wins.README.md,README_CN.md, andskill/SKILL.md(which also now calls out the OpenAI-vs-Anthropic block-shape gotcha).Verification
Dry run of the built binary:
That is byte-for-byte the payload shape I confirmed M3 answers correctly in #224.
Six new tests in
test/commands/text/chat.test.tscover block shape, multi-image, image-without-message, the model override, explicit--modelwinning, and the missing-file error.bun test455 pass / 0 fail;bun run typecheckandbun run lintclean.Need help on this PR? Tag
@codesmith-botwith what you need. Autofix is disabled.