@31337

31337@sh.itjust.works · edit-2 14 hours ago

Larger models train faster (need less compute), for reasons not fully understood. These large models can then be used as teachers to train smaller models more efficiently. I’ve used Qwen 14B (14 billion parameters, quantized to 6-bit integers), and it’s not too much worse than these very large models.

Lately, I’ve been thinking of LLMs as lossy text/idea compression with content-addressable memory. And 10.5GB is pretty good compression for all the “knowledge” they seem to retain.

31337@sh.itjust.works · 2 days ago

My friend just hooks his laptop up to his TV, connects to his VPN, and plays popcorntime (streaming torrents). He used to use streaming sites, but those have been getting taken down left and right.

31337@sh.itjust.works · 4 days ago

Do you remember when it was commonly advised to use fake names and birthdays on online forms, and when “spyware” was a term?