Pitfalls and Opportunities in Google Data Takeout
I've just linked my "Streamlining Google Data Takeout for Google+" guide. This still leaves much of the process manual.
My recommendation to Google is that any manual processing be reduced to the bare minimum possible. This means that either Sane Defaults or URL-based parameters for specification (or both) would help tremendously.
I've made three Google data takeouts so far.
* The first failed entirely.
* The second failed to complete and I didn't download it.
* The third I'm still downloading (it's taken >24h, given failures), and it turns out I'd not specified all relevant components.
* I'm going to have to make a fourth.
Streamlining the process will reduce total load on Google's provisioning systems.
Google could also improve documentation though this is secondary to improving the process itself. Where at all possible, make processes that don't require documentation.
(But don't confuse "undocumented systems" with "documentation not required".)
Full data specs would be really useful, mostly to tool builders for processing and importing to new platforms. It's not clear to me, for example, what "plus_communities" actually contains (and I'm all but certain it doesn't contain what I'd want it to). I suspect data specs are in flux. Docs which update as that spec changes would be hugely useful.
Also the other requests:
* Improved download support.
* Uncapped Google Drive storage.
* Checksums.
* Link-based downloads (wget, curl, rsync, scp, sftp, something). Probably SSH-based.
* A data spec versioning tool.
* Incremental / date-based archive selection.
There's good news about Data Takeout: it's visibly improving and getting attention from Google
Over the past two months people working with Data Takeout (mostly Filip H.F. Slagter and Bernhard Suter, that I'm aware of, both have done truly heroic and useful work) have seen multiple changes to data formats and processes, many specifically requested in Google Feedback. That system isn't all it could be, but for all its limitations it's useful, and there may be better options developing. We strongly encourage anyone using Data Takeout and encountering issues or seeing possible improvements to BOTH submit Feedback and submit posts to this Community. The information is getting to people who can use it.
After nine weeks of radio silence, the official Google+ profile made a post today, and it was about Data Takeout. The official support page for that has been updated and improved. Still not all it could be, but better. Reading the tea leaves is at best fraught, but the place Google are focusing their exceedingly limited public communications on has been Data Takeout. Again, mild encouragement.
If and when there's more information on this, we'll share it.
I've just linked my "Streamlining Google Data Takeout for Google+" guide. This still leaves much of the process manual.
My recommendation to Google is that any manual processing be reduced to the bare minimum possible. This means that either Sane Defaults or URL-based parameters for specification (or both) would help tremendously.
I've made three Google data takeouts so far.
* The first failed entirely.
* The second failed to complete and I didn't download it.
* The third I'm still downloading (it's taken >24h, given failures), and it turns out I'd not specified all relevant components.
* I'm going to have to make a fourth.
Streamlining the process will reduce total load on Google's provisioning systems.
Google could also improve documentation though this is secondary to improving the process itself. Where at all possible, make processes that don't require documentation.
(But don't confuse "undocumented systems" with "documentation not required".)
Full data specs would be really useful, mostly to tool builders for processing and importing to new platforms. It's not clear to me, for example, what "plus_communities" actually contains (and I'm all but certain it doesn't contain what I'd want it to). I suspect data specs are in flux. Docs which update as that spec changes would be hugely useful.
Also the other requests:
* Improved download support.
* Uncapped Google Drive storage.
* Checksums.
* Link-based downloads (wget, curl, rsync, scp, sftp, something). Probably SSH-based.
* A data spec versioning tool.
* Incremental / date-based archive selection.
There's good news about Data Takeout: it's visibly improving and getting attention from Google
Over the past two months people working with Data Takeout (mostly Filip H.F. Slagter and Bernhard Suter, that I'm aware of, both have done truly heroic and useful work) have seen multiple changes to data formats and processes, many specifically requested in Google Feedback. That system isn't all it could be, but for all its limitations it's useful, and there may be better options developing. We strongly encourage anyone using Data Takeout and encountering issues or seeing possible improvements to BOTH submit Feedback and submit posts to this Community. The information is getting to people who can use it.
After nine weeks of radio silence, the official Google+ profile made a post today, and it was about Data Takeout. The official support page for that has been updated and improved. Still not all it could be, but better. Reading the tea leaves is at best fraught, but the place Google are focusing their exceedingly limited public communications on has been Data Takeout. Again, mild encouragement.
If and when there's more information on this, we'll share it.

Using data and electricity for multiple attempts?
ReplyDeleteMy TL;DR is
No, thank you.
#sub
ReplyDeleteDiana Studer My goal is to minimise that, for others. Hence my experimentation.
ReplyDeleteI don't have the best available infrastructure, but my bandwidth is fair, access unmetered, and the takeouts are well below my data caps.
For me it is a benefit of G+ that techie people can and do, while I watch and listen.
ReplyDeleteDiana Studer Bingo.
ReplyDeleteAlmost all of my Data Takeout are GIFs:
ReplyDelete15.775 gif
0.473 png
0.456 json
0.122 jpg
0.010 csv
0.006 webp
0.004 mp4
0.000 ics
What the actual eff?
Considering how much I tend to dislike the damned things, this is ... ironic.
Beautiful! Edward Morbius,
ReplyDelete/following - is it possible to pin this post, please ?
trouble free & Easy backup/download/takeout to be first thing
Thanks
Edward Morbius auto-awesomes perhaps? Videos as gif can quickly use up all that space. Perhaps sort your gifs by file size and look at what the biggest ones actually display?
ReplyDeleteFilip H.F. Slagter There's a lot of them: 5,455. Updated script with counts by filetype (note JSON wins here).
ReplyDelete10,742 0.456 json
8,848 0.010 csv
5,455 15.775 gif
1,291 0.473 png
722 0.122 jpg
79 0.006 webp
56 0.000 ics
4 0.004 mp4
About 2.9 MB each.
And 42.5 kB per JSON file.
Many of the gifs are numbered, e.g., "Takeout/Google+ Stream/Posts/18gsxeu6whw8i.gif(35)" I suspect that's #35 in some animated GIF sequence.
There are also about 160 GIFs above 16 MB in size, with the largest being 18.628 MB. That adds up right there -- 2.56 GB.
On FB I can hide GIFs. On G+ I can at least stop them from dancing at me. But for Takeout - not a chance!
ReplyDeleteI'll wait for the next instalment ;~)