Repository navigation
Interaction between Tensorflow Java and JavaCPP Pointer deallocation #208
Description
Activity
If you're confident you don't need GC, we can set the "org.bytedeco.javacpp.nopointergc" system property to "true" to reduce overhead.
I am not sure wether I need GC or not. Looking at the EagerSession class seems like it is managing memory through the pointer scope. The NDArrays I allocate through DJL/TF also use the same.
I will try the system property you have mentioned and run my benchmarks, thanks for such a quick response!
Reacted by Samuel AudetYou'll need to keep in mind though that neither TF nor DJL are particularly designed to reduce GC. JavaCPP is only a small part of the overall design, and there is a lot garbage that gets generated elsewhere.
Still investigating how GC optimizations work in the DJL + Tensorflow Environment.
There seems to be also a blocking call when using several eager sessions in parallel for multi-threading, and when JavaCPP needs to run GC, it blocks all of the threads with sessions.Like I said, make sure to set the "org.bytedeco.javacpp.nopointergc" system property to "true" to prevent JavaCPP from calling
System.gc(). It should also avoid blocking calls, but let me know if you still see anything of concern there.I've tried the org.bytedeco.javacpp.nopointergc and very quickly ran out of memory :-)
If I keep it "false", then in JVM with G1GC, I spend 25% of the time in GC on a high throughput inference use-case, with occasional blocking.
stacktrace related to allocation
Exception in thread "ForkJoinPool-1-worker-7" java.lang.OutOfMemoryError: Cannot allocate new PointerPointer(1000): totalBytes = 3640M, physicalBytes = 1226M
at org.bytedeco.javacpp.PointerPointer.(PointerPointer.java:149)
at org.tensorflow.EagerOperationBuilder.execute(EagerOperationBuilder.java:270)
at org.tensorflow.EagerOperationBuilder.build(EagerOperationBuilder.java:66)
at org.tensorflow.EagerOperationBuilder.build(EagerOperationBuilder.java:55)
at ai.djl.tensorflow.engine.TfNDArray.(TfNDArray.java:81)
at ai.djl.tensorflow.engine.TfNDManager.create(TfNDManager.java:154)So I've tried it, the blocking calls are removed, the thread with JavaCPP Deallocator is also disappears and no-more blocking GCs. Only issue is I quickly run out of memory :-)
There is a team discussion on DJL around Tensorflow Java and JavaCPP performance, if you are interested - deepjavalibrary/djl#625
Reacted by Samuel AudetSo I've tried it, the blocking calls are removed, the thread with JavaCPP Deallocator is also disappears and no-more blocking GCs. Only issue is I quickly run out of memory :-)
Obviously, when not using the GC, you'll need to make sure to deallocate native memory some other way!
- True that, I've just tested the theory that PointerScope takes care of my specific use case for inference. But you are right, there is quite a lot of garbage besides that.…On Fri, Feb 19, 2021 at 4:33 PM Samuel Audet ***@***.***> wrote: So I've tried it, the blocking calls are removed, the thread with JavaCPP Deallocator is also disappears and no-more blocking GCs. Only issue is I quickly run out of memory :-) Obviously, when not using the GC, you'll need to make sure to deallocate native memory some other way! — You are receiving this because you authored the thread. Reply to this email directly, view it on GitHub <#208 (comment)>, or unsubscribe <https://github.com/notifications/unsubscribe-auth/ABHIJQDEMC63YUQQAQFNHWTS737NBANCNFSM4XD2EC2Q> .Reacted by Samuel Audet and Yunhui YZ Zhang
Eager execution does allocate a lot of native resources, as each operation and each of their outputs will need to be freed.
I don't know much about details the GC feature of JavaCPP we are using now but I can tell that in 1.14, the way it was working is that the GC listener was running in a separate thread and was trying to free resources referenced by deleted objects, but just as a "best effort" since GC is not entirely reliable when it comes to native memory. @saudet, this thread-based implementation was just listening to a phantom reference queue and was therefore non-blocking, is JavaCPP doing something similar?
Ultimately, resources should be cleaned up by closing the eager session enclosing some piece of code and that part is still true:
try (EagerSession s = EagerSession.create()) { Ops tf = Ops.create(s); // your eager computations }This ensure to release all resources independently from the GC. So while it is not enforced by the API, it is recommended to scope your eager operations into multiple eager session instead of only relying on the default one (i.e. the one used when simply invoking
Ops.create()without any parameter).11 remaining items
@karllessard @saudet @skirdey Here is my thought.
There are two strategies of releasing native memory.- Java style: spawn another dedicated thread to clean up native resource object in the ReferenceQueue.
pros: I guess the performance would be slighter better than approach 2 as delete native memory takes time but needs experiments
cons: If the client API invoke the function in asynchronous way (which is usually the case in Java), the dedicated thread is not fast enough to clean up the native resource and cause OOM. Additionally, if engines like pytorch and apache mxnet don't supportthread-safe delete, it will cause the JVM crash. As there might be a small chance that both PointerScope & dedicated thread release the native resource at the same cause and ran into double delete crash. Both PyTorch & Apache MXNet are of this type. So we have no choice but to go with approach 2. Do you know if tensorflow is thread-safe in terms of releasing the native resource like tensor? - C++ style: release the memory as soon as we mark the memory is not used.
pros: The memory consumption is small comparing to approach 1.
cons: The performance could be slightly slower than approach 1 but again it requires experiments.
Either way is trying to get rid of GC (System.gc()) which slows down the entire system. The right pattern of coding Deep Learning system is to make sure all native resources are tracked by a scope and release the memory without a leak and help of GC.
@saudet Does JavaCpp support approach 2 if users are working on memory-intensive application?
- Java style: spawn another dedicated thread to clean up native resource object in the ReferenceQueue.
It stores a strong reference, otherwise this kind of pattern may result in a crash:
Ok, this new paradigm introduced once we have switched to JavaCPP is probably the cause of the issue here. Before, the
EagerSessionscope was acting weakly and would not prevent the native memory of its garbage-collected resources to be released.Some sessions (like the default one, which I think @skirdey uses) are never or rarely closed, therefore entirely rely on the GC or on explicit closing of the resources themselves (the tensors, for example). Since a strong reference is kept on the native pointer, the memory will never be released. That certainly happens with the
EagerOperationresources.It could be great if a user can decide to attach weakly or strongly a
Pointerto aPointerScope. If that's not feasible, then we might want to prevent attaching these resources to the eager session if we know that session is the default one. But there is probably a cleaner way to fix this as well.@skirdey , do you build your own version of TF Java when you are doing your tests or you use the prebuilt one coming with DJL? It would be interesting to see the behaviour of your code by adding this patch.
Looking at DJL code and TfNDManager it does create a single session in with async option https://github.com/awslabs/djl/blob/master/tensorflow/tensorflow-engine/src/main/java/ai/djl/tensorflow/engine/TfNDManager.java#L69
I can not see if it is a "default" session or not.
It stores a strong reference, otherwise this kind of pattern may result in a crash:
Ok, this new paradigm introduced once we have switched to JavaCPP is probably the cause of the issue here. Before, the
EagerSessionscope was acting weakly and would not prevent the native memory of its garbage-collected resources to be released.Some sessions (like the default one, which I think @skirdey uses) are never or rarely closed, therefore entirely rely on the GC or on explicit closing of the resources themselves (the tensors, for example). Since a strong reference is kept on the native pointer, the memory will never be released. That certainly happens with the
EagerOperationresources.I see, that's what @rnett was referring in pull #188 (comment).
It could be great if a user can decide to attach weakly or strongly a
Pointerto aPointerScope. If that's not feasible, then we might want to prevent attaching these resources to the eager session if we know that session is the default one. But there is probably a cleaner way to fix this as well.It's never a good idea to rely on the GC. We cannot avoid situations like I mention above where the native side keeps a reference to an existing Tensor, but where the Java side doesn't. The only sane way to deal with this is by not relying on the GC at all. We can still have it as a sort of option for users that don't want to use something like TensorScope though, but in that case, since we cannot offer any guarantees anyway, it makes no difference whether eager session holds on to weak references, or no reference at all!
@skirdey , do you build your own version of TF Java when you are doing your tests or you use the prebuilt one coming with DJL? It would be interesting to see the behaviour of your code by adding this patch.
Which patch?
It stores a strong reference, otherwise this kind of pattern may result in a crash:
Ok, this new paradigm introduced once we have switched to JavaCPP is probably the cause of the issue here. Before, the
EagerSessionscope was acting weakly and would not prevent the native memory of its garbage-collected resources to be released.
Some sessions (like the default one, which I think @skirdey uses) are never or rarely closed, therefore entirely rely on the GC or on explicit closing of the resources themselves (the tensors, for example). Since a strong reference is kept on the native pointer, the memory will never be released. That certainly happens with theEagerOperationresources.
It could be great if a user can decide to attach weakly or strongly aPointerto aPointerScope. If that's not feasible, then we might want to prevent attaching these resources to the eager session if we know that session is the default one. But there is probably a cleaner way to fix this as well.
@skirdey , do you build your own version of TF Java when you are doing your tests or you use the prebuilt one coming with DJL? It would be interesting to see the behaviour of your code by adding this patch.Looking at DJL code and TfNDManager it does create a single session in with async option https://github.com/awslabs/djl/blob/master/tensorflow/tensorflow-engine/src/main/java/ai/djl/tensorflow/engine/TfNDManager.java#L69
I can not see if it is a "default" session or not.
I checked the source code. It should not be a default EagerSession. The defaultEagerSession is created only when you call initDefault(options) or getDefault()
@karllessard @saudet @skirdey Here is my thought.
There are two strategies of releasing native memory.-
Java style: spawn another dedicated thread to clean up native resource object in the ReferenceQueue.
pros: I guess the performance would be slighter better than approach 2 as delete native memory takes time but needs experiments
cons: If the client API invoke the function in asynchronous way (which is usually the case in Java), the dedicated thread is not fast enough to clean up the native resource and cause OOM. Additionally, if engines like pytorch and apache mxnet don't supportthread-safe delete, it will cause the JVM crash. As there might be a small chance that both PointerScope & dedicated thread release the native resource at the same cause and ran into double delete crash. Both PyTorch & Apache MXNet are of this type. So we have no choice but to go with approach 2. Do you know if tensorflow is thread-safe in terms of releasing the native resource like tensor? -
C++ style: release the memory as soon as we mark the memory is not used.
pros: The memory consumption is small comparing to approach 1.
cons: The performance could be slightly slower than approach 1 but again it requires experiments.
Either way is trying to get rid of GC (System.gc()) which slows down the entire system. The right pattern of coding Deep Learning system is to make sure all native resources are tracked by a scope and release the memory without a leak and help of GC.
@saudet Does JavaCpp support approach 2 if users are working on memory-intensive application?
JavaCPP supports both the "Java style", that is using GC via
PhantomReferenceand aReferenceQueue, and "C++ style" viaPointerScope. Currently, JavaCPP only uses 1 "deallocator thread", but it is possible to extend that to support multiple threads and increase throughput. However, there are too many limitations to that approach using GC that I don't think it's worth spending too much time on that. Like you mention, it's problematic when the native allocator isn't thread-safe, but even when it is thread-safe, allocating more memory than required leads to memory fragmentation and poor performance, typically poorer than managing it "C++ style" in my experience. So it is almost always a better idea to manage native memory like one would do it in C++ or Python. That's exactly the paradigm thatPointerScopeattempts to bring to Java, that we can use across libraries, like I showed here for the C++ APIs of TensorFlow and OpenCV: http://bytedeco.org/news/2018/07/17/bytedeco-as-distribution/-
@karllessard I'm trying to find where in the old code it was using weak references, and I can't find. I don't remember removing anything like that myself either. From what I can tell, EagerSession has always been keeping strong references with this map: https://github.com/tensorflow/tensorflow/blob/master/tensorflow/java/src/main/java/org/tensorflow/EagerSession.java#L499
NativeReferenceis a subclass ofjava.lang.ref.PhantomReference.I checked the source code. It should not be a default EagerSession. The defaultEagerSession is created only when you call initDefault(options) or getDefault()
@stu1130 , even if it is not the default session, if the created session remains open for a long time while creating a lot of ops and tensors, the issue can happen
@karllessard I'm trying to find where in the old code it was using weak references, and I can't find. I don't remember removing anything like that myself either. From what I can tell, EagerSession has always been keeping strong references with this map
@saudet: like @Craigacp pointed out,
NativeReferenceextendsPhantomReferenceand are extended by each type of native resource, like here. So basically how it works is that if aNativeReferencewas queued by the GC, itsdelete()method was invoked to release the native resources and it was removed from session the session scope. If it was never queued but the session was closed, all remainingNativeReferencein the session scope were deleted explicitly. So what I was saying is that the eager session was holding references acting "weakly" and not necessarily stackingWeakReferenceobjects.Which patch?
There is none yet :) Ideally, like we discussed, it would be possible to refer weakly to a
Pointerin aPointerScopefrom JavaCPP directly. Otherwise we have to find another way to restore the previous/correct behavior. I was curious to know if @skirdey was already setup to try a patch if we have one but I guess I can easily reproduce the problem on my side and validate this hypothesis.NativeReferenceis a subclass ofjava.lang.ref.PhantomReference.Ah, that's where I screwed up. I remember getting confused about what
NativeReferencewas exactly and associating it withPointerinstead of more likePhantomReference<Pointer>. Well, let's see, that's a pattern currently supported by neither theCleanerin JDK 9+ nor theResourceScopeas proposed by Panama, so it's probably not something I want to encourage users to start doing as part of JavaCPP either... See the point I make in pull #188 (comment).If we still want to have something that behaves like the original implementation, we should probably just replace the
PointerScopefield inEagerSessionwith some sort of collection ofWeakReference<Pointer>, as currently proposed by @rnett forTensorScopeas per pull #188 (review).Ok, I understand that you want JavaCPP to follow closely the (future) behaviour of the JDK.
I'm personally totally comfortable with what @rnett proposed about weak tensor scopes, which is quite identical to the original behaviour of the eager sessions. So TF could have its own "weak" pointer scope. @rnett , were you planning to change your proposed
TensorScopeto make it more generic for any kind of native resources, like the eager ops?If that's reshuffling too much of your work, we can simply keep our own weak references directly in the
EagerSessionlike before, and since JavaCPP will still take care of the cleaner portion, it should really just be a small change to do.I wasn't planning on doing it as part of TensorScope, but rather as part of Ops or Scope since we're passing it around anyways. But that's all up for discussion, I don't have anything particularly firm yet.
Ok let's play it simple for now and I'll try to see how it works by simply replacing the
PointerScopeinEagerSessionby a list of weak references, I'll let you all know of the results.Ok, so the issue was very easy to reproduce. Starting a JVM with only 256M of memory, this simple loop was hitting a OOM between 30K and 40K iterations:
public static void main(String[] args) { try (EagerSession s = EagerSession.create()) { Ops tf = Ops.create(s); while (true) { tf.math.add(tf.constant(2), tf.constant(2)); } } }
I pushed this PR #229 to keep only weak references on eager resources in the session, as proposed earlier, and now the garbage collection allows this loop to run forever. I'm pretty confident this should fix the issue observed earlier by @skirdey when using DJL.
- Awesome, I can give a try!…On Sunday, 28 February 2021, Karl Lessard ***@***.***> wrote: Ok, so the issue was very easy to reproduce. Starting a JVM with only 256M of memory, this simple loop was hitting a OOM between 30K and 40K iterations: public static void main(String[] args) { try (EagerSession s = EagerSession.create()) { Ops tf = Ops.create(s); while (true) { tf.math.add(tf.constant(2), tf.constant(2)); } } } I pushed this PR to keep only weak references on eager resources in the session, as proposed earlier, and now the garbage collection allows this loop to run forever. I'm pretty confident this should fix the issue observed earlier by @skirdey <https://github.com/skirdey> when using DJL. — You are receiving this because you were mentioned. Reply to this email directly, view it on GitHub <#208 (comment)>, or unsubscribe <https://github.com/notifications/unsubscribe-auth/ABHIJQGJ5QAGOTI5ZNW3WW3TBMGENANCNFSM4XD2EC2Q> .Reacted by Karl Lessard and Jake Lee
- added a commit that references this issue
on Mar 2, 2021

I am trying to understand why when I use DJL + Tensorflow engine, there is a good amount of time spent in GC, while both DJL and Tensorflow Java seem to use pointerscope and do not relay on GC for object cleanup.
For example, I see JavaCPP Pointer.deallocator gets invoked https://github.com/bytedeco/javacpp/blob/master/src/main/java/org/bytedeco/javacpp/Pointer.java#L666-L667
which is a heavy synchronized call.
Any help is appreciated.