I don't know about tensor flow in particular but are little-known methods of running "general purpose" parallel programs on GPUs. Specifically, H. Dietz' MOG, "Mimd on GPU". It's a shame the project hasn't gotten more attention imo.
See: https://en.wikipedia.org/wiki/Flynn%27s_taxonomy for explanations of terms.
Sorry for my late reply, I just wanted to thank you because your comment is an exceptionally good example of what I was trying to get at with my longwinded explanations. Compiling MIMD to SIMD is the future of programming, although it seems that companies will try every other course of action before they realize that.