The instinct is to upload the photo and ask for a finished graphic in one shot. That is also how you end up with a beautiful infographic that explains something incorrectly, because the model misread a word in the image and you never saw the mistake.
Stage it instead, in an image-capable chat tool like ChatGPT. Upload a legible photo and ask one narrow question at a time.
The point of the order is the gap between steps 2 and 3. You read the first two answers and catch a misreading of the source while it is still a sentence you can correct, rather than after it has been set in a polished graphic you are about to share.
A good-looking infographic carries more authority with a reader than a paragraph does, so it deserves more verification, not less.
If the graphic needs Hebrew in it, see How do I keep AI image tools from garbling the Hebrew on my flyers?.