Summary

I would like to propose evolving ChatGPT from a collection of separate capabilities into one persistent, permission-based personal assistant that can operate naturally across a user’s devices.

The goal is not unrestricted device access. The goal is controlled agency: ChatGPT should be able to perceive, remember and act only within transparent boundaries explicitly defined by the user.

1. Granular Delegated Agency

Users should be able to grant specific, independently revocable and optionally time-limited permissions for email, calendar, files, notifications, microphone, camera, screen, location and connected applications.

Permissions should distinguish between different actions. For example, a user might allow ChatGPT to read and organize email while prohibiting deletion or sending.

High-impact actions could still require explicit confirmation.

Examples:

  • “Organize my banking emails, but do not delete anything.”
  • “Access this folder for the next hour, but do not share any files.”
  • “Watch my screen while I configure this device and guide me.”

2. Cross-Device Continuity / Distributed Embodiment

A phone, smartwatch, earbuds, computer and future wearable or robotic devices should not behave as separate assistants.

They should function as different interfaces, sensors and action points for the same persistent assistant.

Memory, identity, preferences and ongoing context should belong to the assistant rather than to one particular device.

Phone + smartwatch + computer + earbuds should not mean four assistants. They should be four interfaces to the same assistant.

Each device would contribute only the capabilities explicitly permitted by the user.

3. Contextual Prospective Memory

The assistant should remember future intentions expressed naturally in conversation and surface them when the appropriate real-world context occurs.

For example, at home I might casually say:

“There’s no cheese left. Let’s buy some when we go out.”

I should not need to create a reminder, select a time, open an application or type anything.

Later, if we pass a supermarket and the necessary permissions are enabled, ChatGPT could simply say:

“Should we buy the cheese here?”

This is different from a conventional reminder.

The assistant understood an intention, retained it, recognized the appropriate real-world context and surfaced it when it became useful.

4. Human Effort Minimization

Computers can perform trillions of operations, yet humans still spend enormous amounts of time adapting themselves to computers: navigating menus, typing commands, correcting text, switching applications and repeatedly providing information that a sufficiently contextual system could already understand.

AI should reverse this relationship.

The human should express intent in the most natural available way; the system should bear the burden of translating that intent into digital operations.

The success metric should therefore not simply be how many features an AI assistant provides, but how much unnecessary human interaction those capabilities eliminate.

The Larger Idea

I am not proposing another collection of independent AI features.

I am proposing that existing and future capabilities such as memory, voice, camera, screen understanding, connected applications, notifications, location and device sensors should be able to work as components of one continuous personal assistant.

The assistant should maintain continuity while the user moves between devices, remember unfinished intentions and act only within permissions established by the user.

Privacy, transparency and reversibility should be fundamental to the architecture.

One assistant. Multiple devices. Persistent context. User-defined boundaries. Minimum human effort.

The objective is simple: technology should adapt to the human, rather than requiring the human to continually adapt to technology.