interleaved_multimodal
`<image> 텍스트 <image> 텍스트 …` 처럼 이미지와 텍스트를 한 시퀀스에 자유롭게 섞은 입력. GPT-3식 in-context 퓨샷을 멀티모달로 확장한다.