google cloud professional machine learning engineer

/

google cloud professional machine learning engineer認定

タグ


参考ドキュメント

tensorflowの分散トレーニング

  • TPU > GPU > CPU
  • synronouse training
    • different slices of input data in sync
    • aggregationg gradients at each step
  • async training
    • all worker are independent
  • mirrored strategy
    • sync
    • GPUs on one machine
    • creates one replica per GPU
    • all variable are mirrored across all replicas
  • NVIDIA NCCL
    • all-reduce implementation
  • parameter server strategy
    • parameter server
  • central storage strategy
    • sync
    • not mirred
  • tpu strategy
    • mirred
    • all-reduce

TPUでのトレーニング

ai platformでの分散トレーニング

  • master, worker, paramter serverのcontainerでトレーニングする
  • sklearn, xgboost等はサポートされない
  • ベイス最適化がある

説明可能なAI

  • アルゴリズム
    • Integrated Gradients
    • XAI
    • Sampled Shaplely
  • AutoML tableもサポート

kubeflowパイプライン

精度

  • 再現率
  • 適合性
  • F1
  • AUCROC > AUCRP(クラスバランス不変)