Scalaz 7 zipWithIndex / group 열거로 메모리 누수 방지

program story

Scalaz 7 zipWithIndex / group 열거로 메모리 누수 방지

inputbox 2020. 8. 10. 08:02

Scalaz 7 zipWithIndex / group 열거로 메모리 누수 방지

배경

이 질문 에서 언급했듯이 Scalaz 7 반복을 사용하여 일정한 힙 공간에서 대규모 (즉, 제한되지 않은) 데이터 스트림을 처리하고 있습니다.

내 코드는 다음과 같습니다.

type ErrorOrT[M[+_], A] = EitherT[M, Throwable, A]
type ErrorOr[A] = ErrorOrT[IO, A]

def processChunk(c: Chunk, idx: Long): Result

def process(data: EnumeratorT[Chunk, ErrorOr]): IterateeT[Vector[(Chunk, Long)], ErrorOr, Vector[Result]] =
  Iteratee.fold[Vector[(Chunk, Long)], ErrorOr, Vector[Result]](Nil) { (rs, vs) =>
    rs ++ vs map { 
      case (c, i) => processChunk(c, i) 
    }
  } &= (data.zipWithIndex mapE Iteratee.group(P))

문제

메모리 누수가 발생한 것 같지만 버그가 Scalaz에 있는지 내 코드에 있는지 알 수있을만큼 Scalaz / FP에 익숙하지 않습니다. 직관적으로이 코드 는 -size 공간의 P 배 ( Chunk대략 ) 만 필요 합니다.

참고 :에서 발생한 유사한 질문 을 찾았 OutOfMemoryError지만 내 코드는 consume.

테스팅

문제를 격리하기 위해 몇 가지 테스트를 실행했습니다. 요약하면, 누출시에만 모두 발생하는 표시 zipWithIndex및 group사용됩니다.

// no zipping/grouping
scala> (i1 &= enumArrs(1 << 25, 128)).run.unsafePerformIO
res47: Long = 4294967296

// grouping only
scala> (i2 &= (enumArrs(1 << 25, 128) mapE Iteratee.group(4))).run.unsafePerformIO
res49: Long = 4294967296

// zipping and grouping
scala> (i3 &= (enumArrs(1 << 25, 128).zipWithIndex mapE Iteratee.group(4))).run.unsafePerformIO
java.lang.OutOfMemoryError: Java heap space

// zipping only
scala> (i4 &= (enumArrs(1 << 25, 128).zipWithIndex)).run.unsafePerformIO
res51: Long = 4294967296

// no zipping/grouping, larger arrays
scala> (i1 &= enumArrs(1 << 27, 128)).run.unsafePerformIO
res53: Long = 17179869184

// zipping only, larger arrays
scala> (i4 &= (enumArrs(1 << 27, 128).zipWithIndex)).run.unsafePerformIO
res54: Long = 17179869184

테스트 용 코드 :

import scalaz.iteratee._, scalaz.effect.IO, scalaz.std.vector._

// define an enumerator that produces a stream of new, zero-filled arrays
def enumArrs(sz: Int, n: Int) = 
  Iteratee.enumIterator[Array[Int], IO](
    Iterator.continually(Array.fill(sz)(0)).take(n))

// define an iteratee that consumes a stream of arrays 
// and computes its length
val i1 = Iteratee.fold[Array[Int], IO, Long](0) { 
  (c, a) => c + a.length 
}

// define an iteratee that consumes a grouped stream of arrays 
// and computes its length
val i2 = Iteratee.fold[Vector[Array[Int]], IO, Long](0) { 
  (c, as) => c + as.map(_.length).sum 
}

// define an iteratee that consumes a grouped/zipped stream of arrays
// and computes its length
val i3 = Iteratee.fold[Vector[(Array[Int], Long)], IO, Long](0) {
  (c, vs) => c + vs.map(_._1.length).sum
}

// define an iteratee that consumes a zipped stream of arrays
// and computes its length
val i4 = Iteratee.fold[(Array[Int], Long), IO, Long](0) {
  (c, v) => c + v._1.length
}

질문

내 코드에 버그가 있습니까?
이 작업을 일정한 힙 공간에서 어떻게 만들 수 있습니까?

이것은 이전 iterateeAPI를 고수하는 사람에게는 별 도움이되지 않을 것이지만 최근에 동일한 테스트가 scalaz-stream API 에 대해 통과 함을 확인했습니다 . 을 (를) 대체하기위한 최신 스트림 처리 API입니다 iteratee.

완전성을 위해 다음은 테스트 코드입니다.

// create a stream containing `n` arrays with `sz` Ints in each one
def streamArrs(sz: Int, n: Int): Process[Task, Array[Int]] =
  (Process emit Array.fill(sz)(0)).repeat take n

(streamArrs(1 << 25, 1 << 14).zipWithIndex 
      pipe process1.chunk(4) 
      pipe process1.fold(0L) {
    (c, vs) => c + vs.map(_._1.length.toLong).sum
  }).runLast.run

이것은 n매개 변수에 대한 모든 값과 함께 작동합니다 (충분히 기다릴 수있는 경우)-2 ^ 14 32MiB 배열 (즉, 시간이 지남에 따라 총 절반 TiB의 메모리가 할당 됨)으로 테스트했습니다.

참고 URL : https://stackoverflow.com/questions/19128856/avoiding-memory-leaks-with-scalaz-7-zipwithindex-group-enumeratees

'program story' 카테고리의 다른 글

vi 다시 그리기 화면을 만드는 방법은 무엇입니까? (0)	2020.08.10
Swift에서 UIColorFromRGB를 어떻게 사용할 수 있습니까? (0)	2020.08.10
나침반을 포함한 Android 휴대 전화 방향 개요 (0)	2020.08.10
XPath의 인덱스가 0이 아닌 1로 시작하는 이유는 무엇입니까? (0)	2020.08.10
JVM이 JIT 컴파일 된 코드를 캐시하지 않는 이유는 무엇입니까? (0)	2020.08.10

현재글Scalaz 7 zipWithIndex / group 열거로 메모리 누수 방지

inputbox

Scalaz 7 zipWithIndex / group 열거로 메모리 누수 방지

Scalaz 7 zipWithIndex / group 열거로 메모리 누수 방지

'program story' 카테고리의 다른 글

'program story'의 다른글

티스토리툴바

Scalaz 7 zipWithIndex / group 열거로 메모리 누수 방지

Scalaz 7 zipWithIndex / group 열거로 메모리 누수 방지

'program story' 카테고리의 다른 글

'program story'의 다른글

관련글

티스토리툴바