mediumAI Engineering

How do multimodal models handle images, audio and other non-text inputs?

871 views
01

Understand the problem

Vision encoders, audio tokens, and what multimodality means for app design.

multimodalvisionaudiofundamentals
02

Attempt it yourself

Sketch your approach before reading the solution — that's what interviews test.

Stuck? AI Nudge Available

Get a conceptual hint to guide your logic without spoiling the final implementation.

03

Study the solution

The solution is waiting

Give it an honest attempt first — then compare your thinking with the full walkthrough.

04

Read the code

Vision request with cost control
const resized = await sharp(buf).resize({ width: 1024 }).toBuffer(); // cap tokens

await client.messages.create({
  model,
  max_tokens: 500,
  messages: [{
    role: "user",
    content: [
      { type: "image", source: { type: "base64", media_type: "image/png",
        data: resized.toString("base64") } },
      { type: "text", text: "From this invoice, extract vendor, date, total as JSON." },
    ],
  }],
});
05

Join the discussion

Discussion (0)

Sign in to join the discussion.

No responses yet. Be the first to share what you think.