To get good context, you need video data. To get good video data, you need cameras, lots of them, on your head.
To get them on your head you need them to be small and light
To get them to run you need batteries. But the maximum battery size you can get, without wires is about 1-1.5 watt hour.
Then, without making a custom wireless protocol, the maximum bandwith you can reliabily expect to user (with an iphone) is 1megabit.
That means you need to compress the world around you to 1megatbit a second.
Now, with a bit or work, like using eye tracking to segment what you are actually looking at(usually a 64x64 pixel image at 15-20 frames a second), and occasional wide angle view when the scene changed, you can build really good context, even reading a book.
But, you also need good location data, along with room classfication to get context. You can't really use GPS because they aren't accurate enough, and don't work all that well indoors. So you use visual odometry.
Once you have all that, you then now need to teach the machine to understand that 1mbit stream to work out if it knows the answer you need. oh and that has to work in a device that has ~10-15 watthours. or offload to a bit boy machine over a patching network.
Aar is a welcome enhancement (that has its own, but potentially less privacy issues than video capture). There's a whole set of aar enabled glasses that focus more on the user via audio rather than reckless complete video capture like so many "smart" pervert glasses being shilled by large advertising companies.
> They deliver high-quality sound and give developers access to a range of controls and sensor data, from volume to head movement to biometrics. Built-in microphones and simple controls make it easy to trigger assistants and audio apps.
To get good context, you need video data. To get good video data, you need cameras, lots of them, on your head.
To get them on your head you need them to be small and light
To get them to run you need batteries. But the maximum battery size you can get, without wires is about 1-1.5 watt hour.
Then, without making a custom wireless protocol, the maximum bandwith you can reliabily expect to user (with an iphone) is 1megabit.
That means you need to compress the world around you to 1megatbit a second.
Now, with a bit or work, like using eye tracking to segment what you are actually looking at(usually a 64x64 pixel image at 15-20 frames a second), and occasional wide angle view when the scene changed, you can build really good context, even reading a book.
But, you also need good location data, along with room classfication to get context. You can't really use GPS because they aren't accurate enough, and don't work all that well indoors. So you use visual odometry.
Once you have all that, you then now need to teach the machine to understand that 1mbit stream to work out if it knows the answer you need. oh and that has to work in a device that has ~10-15 watthours. or offload to a bit boy machine over a patching network.
ie: https://www.projectaria.com/
Fear the Greeks when they give you presents.
But i fear it is too late.