Giving images, screenshots and files to the AI
The AI can read more than text. How to upload photos, screenshots and PDFs, what comes out of it and where the limits are.
Most people type everything out for the AI. They laboriously copy the text from a screenshot into the chat, even though they could simply upload the image. It is the most overlooked beginner trick of 2026. ChatGPT, Claude and Gemini do not just read text. They look at photos, read screenshots, work their way through entire PDFs. You only have to know where the button is and how to ask the question.
Where you upload
In each of the three big tools there is a paperclip or plus symbol in the input field. You click on it, pick a file from your computer or phone, and write your question along with it. On a phone you can usually open the camera directly and take a photo. On a computer you can often just drag the file into the chat window. That was the whole magic.
The second part is the important one: the file alone is not enough. You have to say what you want. A photo without a question gets you a general description. A photo with a clear question gets you the answer you need.
Photos and screenshots
Here are a few things that save time immediately in everyday life.
A screenshot of an error message you do not understand. Upload it and ask "What does this mean and what should I do?" The AI reads the message and explains it in plain English.
A photo of a handwritten note or of a whiteboard after a meeting. Ask "Type this up cleanly and turn it into a bullet list." Scrawly handwriting becomes a tidy text.
A screenshot of a table or a chart. Ask "What are the three most important numbers here?" or "Explain to me what this graphic is saying." It also works with diagrams that you cannot make sense of on the fly.
A product photo or a picture of a plant, a device, a component. Ask "What is this and how do I use it?" The recognition is good, but not perfect. More on that in a moment.
PDFs and long documents
For job seekers and freelancers this is the biggest lever. You upload a PDF and let it be worked through, instead of reading twenty pages yourself.
A rental contract, an employment contract, a letter from your insurer. Ask "Summarize the most important points and flag everything that could be a disadvantage for me." In thirty seconds you get an overview that would otherwise take you half an hour to work out.
A job posting as a PDF. Upload it together with your CV and ask "Where do I fit well, where am I missing something, and which three sentences from my CV should I rewrite for this position?" Two files at once, one concrete result.
A specialist article or a manual. Ask "Explain the core topic to me as if I were fifteen" or "What do I really need to know from this?" Good for reading your way into a new subject.
What works well and what does not
Works well: printed text, clear screenshots, tables, charts, ordinary photos of objects. Tidy handwriting is read reliably too.
It gets harder with bad photos. A picture taken at an angle, blurred or dark leads to misreadings. With very scrawly handwriting the AI sometimes guesses instead of reading. And with numbers from a tight table you get the most errors, because a digit slips out of place easily.
Hence the iron rule: with anything where a wrong number gets expensive, so invoices, contracts, appointments, you check the result against the original. The AI is a fast first look here, not the final authority. That applies just as much as it does with text it invents. How to check that systematically is in the routine for the fact check.
How to phrase the question about the image
Three things make the difference.
Say what the file is. "This is a screenshot from my online banking" gives the AI context that it cannot reliably get out of the image alone.
Say what you want. Not "what is this", but "read me the amounts as a list" or "explain the second paragraph to me". The more concrete the task, the more usable the answer.
Ask back when something is unclear. "Are you sure about the third number, or is the image too blurry there?" A good model will then tell you honestly where it guessed.
Briefly on data protection
A photo of a document often contains more than you think. Names, addresses, account numbers, maybe other notes in the background. Before you upload sensitive things, the same applies as with text: consider whether the data really belongs in somebody else's service, and switch off the training use if you have not done that yet. The details are in the lesson on data protection for beginners.
Next step
Try it out once today. Take a photo of anything you cannot get sorted right now, a note, a table, an error message, and upload it with a clear question attached. Once you have seen how quickly an image turns into usable text, you will never type everything out yourself again.