Just like input and output modalities, please also include the tool result modality, because, for example, Gemini models support file,image, video, and audio as tool results, while the Anthropic models support text and image, and the Openai models support text only.
Just like input and output modalities, please also include the tool result modality, because, for example, Gemini models support file,image, video, and audio as tool results, while the Anthropic models support text and image, and the Openai models support text only.