OceSha VenturesOceSha Ventures
Back to Insights

AI Strategy

Why Are We Still Typing to Computers?

Humans speak, point, show and change their minds mid-thought. Software mostly still asks us to type. That gap is finally starting to close.

By Rohan HallAI Technologist, Author & EducatorLinkedIn
September 3, 2026 · 3 min read
Hands near a keyboard next to a microphone icon, representing a choice of input

Think about how you actually explain something complicated to another person. You speak. You point at a screen. You show them a photo or a document. You get halfway through and change your mind about what you're actually asking. You add context as it occurs to you, not in a fixed order. None of that resembles filling out a form or typing a search query, yet typing is still how most software expects us to communicate.

Graphical interfaces were a huge improvement, not the end state

The move from command lines to graphical interfaces made computing accessible to an enormous number of people who would never have learned a command syntax. Icons, menus and clickable buttons were a genuine leap forward. But they did not remove the underlying requirement: a person still had to learn the machine's model of the task. Where is the setting. What is this feature called. Which menu holds the option I need. The interface got friendlier, not fluent.

Typing extended that same pattern into the web. Search boxes and forms are still asking a person to translate what they want into the system's expected shape: the right keywords, the right field, the right format. It is a skill, and most of us are so practiced at it we no longer notice we are performing one.

Another option, not a replacement

Voice and multimodal AI introduce something genuinely different: the option to just describe what you're trying to accomplish, in your own words, and let the system do the work of interpreting that. "I'm trying to figure out if this plan covers my situation" is a perfectly reasonable thing to say to a person. Increasingly, it's a reasonable thing to say to an intelligent interface as well.

This is not an argument that typing goes away. It doesn't, and it shouldn't. The honest claim is narrower and more useful: the future is not voice instead of text, it's the person choosing whichever mode is most natural for the moment they're in.

Context decides the mode

  • Open office or shared space — typing or reading is often more appropriate than speaking aloud.
  • Noisy environment — text is more reliable than voice recognition fighting background sound.
  • Privacy-sensitive topics — a person may prefer to type something they wouldn't want overheard.
  • Precision tasks — entering an exact account number, address or code is usually faster and more accurate typed.
  • Mobile or hands-busy situations — speaking while walking or driving is far more natural than typing.
  • Accessibility needs — voice can remove barriers for people who find typing difficult, and text can do the same for people who find speech difficult.

None of these situations is universal, and none of them is permanent even for the same person — the same visitor might type at their desk and speak from their phone an hour later. A system that only offers one mode is quietly excluding whichever group that mode doesn't suit.

What this means for leaders building or buying digital experiences

The practical question is not "should we add voice." It's "where does each mode reduce friction for the people we serve, and are we building something flexible enough to let them decide." A support flow that only accepts typed queries is asking every visitor to be equally comfortable with text, which is never true across a real audience. A sales inquiry that only accepts voice excludes people who'd rather not be overheard, or who are in a meeting.

The more durable answer is an interface layer that treats voice, text, images and documents as equally valid ways in, understands intent regardless of which one was used, and lets the person switch mid-conversation without losing context. That is a meaningfully different design goal than "add a chatbot" or "add a voice button," and it changes how the underlying system needs to be built, not just how it looks on the surface. It also connects directly to the multilingual dimension of intelligent interfaces — mode and language are two separate forms of the same underlying idea: let the person communicate naturally.

The keyboard isn't going anywhere

Typing remains, in many situations, the fastest and most precise tool available, and it will stay part of how people interact with software for a long time. The goal isn't to retire it. It's to stop treating it as the only acceptable way in, when it was really just the easiest one for the machine.

Curious what a genuinely multimodal experience feels like in practice? Visit ocesha.com and try speaking to it instead of typing.

Work with OceSha Ventures

Have an AI Opportunity You Want to Build?

OceSha Ventures helps organizations turn AI opportunities into working software, automation, agents, and applications.

Discuss Your Project